Ana içeriğe geç
google avatarı

Gemini 3.1 Flash TTS API

google/gemini-3.1-flash-tts

Metni 30 sesle ifadeli ve yönlendirilebilir bir konuşmaya dönüştürün — okuyuşu düz bir dille tarif edin, 24 kHz WAV alın.

üretilen ses için saniye başına 0.0006

Model girdisi

Input

Style direction plus the words to speak, as '{style instruction}: {text}' — e.g. 'Say the following in a warm, curious way: OK, so... tell me about this AI thing.' Direction controls accent, pace, tone and emotion. Inline audio tags such as [laughs] or [sigh] are supported. Combined direction and text must be under ~8,000 bytes; audio beyond ~655 seconds is truncated.

Prebuilt voice to speak in. Each has a distinct character, e.g. Kore (firm), Puck (upbeat), Aoede (breezy), Charon (informative), Sulafat (warm), Enceladus (breathy).

Locale for the delivery. Any prebuilt voice can be paired with any of these locales.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Loading
Generated in 4.574 seconds
Logs (1 lines)

Örnek istekler

Örnekler

Gemini 3.1 Flash TTS API

Gemini 3.1 Flash TTS is a sound AI model by google. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.0006 per second of audio.

POST https://queue.modelrunner.run/google/gemini-3.1-flash-tts

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/google/gemini-3.1-flash-tts \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "voice": "Vindemiatrix",
    "prompt": "Sakin ve güven veren bir anlatıcı gibi oku: Sabahın ilk ışığıyla birlikte liman yavaşça uyanıyordu.",
    "language_code": "tr-tr",
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/google/gemini-3.1-flash-tts/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/google/gemini-3.1-flash-tts/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("google/gemini-3.1-flash-tts", {
  input: {
    "voice": "Vindemiatrix",
    "prompt": "Sakin ve güven veren bir anlatıcı gibi oku: Sabahın ilk ışığıyla birlikte liman yavaşça uyanıyordu.",
    "language_code": "tr-tr"
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/google/gemini-3.1-flash-tts",
    headers=headers,
    json={
      "voice": "Vindemiatrix",
      "prompt": "Sakin ve güven veren bir anlatıcı gibi oku: Sabahın ilk ışığıyla birlikte liman yavaşça uyanıyordu.",
      "language_code": "tr-tr"
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of Gemini 3.1 Flash TTS
NameTypeRequiredDescription
promptstringyesStyle direction plus the words to speak, as '{style instruction}: {text}' — e.g. 'Say the following in a warm, curious way: OK, so... tell me about this AI thing.' Direction controls accent, pace, tone and emotion. Inline audio tags such as [laughs] or [sigh] are supported. Combined direction and text must be under ~8,000 bytes; audio beyond ~655 seconds is truncated.
voiceenumnoPrebuilt voice to speak in. Each has a distinct character, e.g. Kore (firm), Puck (upbeat), Aoede (breezy), Charon (informative), Sulafat (warm), Enceladus (breathy). One of: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Puck, Pulcherrima, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr, Zubenelgenubi. Default: "Kore".
language_codeenumnoLocale for the delivery. Any prebuilt voice can be paired with any of these locales. One of: ar-eg, bn-bd, de-de, en-in, en-us, es-es, fr-fr, hi-in, id-id, it-it, ja-jp, ko-kr, mr-in, nl-nl, pl-pl, pt-br, ro-ro, ru-ru, ta-in, te-in, th-th, tr-tr, uk-ua, vi-vn. Default: "en-us".

Machine-readable: OpenAPI schema · llms.txt

Use Gemini 3.1 Flash TTS from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and Gemini 3.1 Flash TTS becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint google/gemini-3.1-flash-tts.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run google/gemini-3.1-flash-tts on ModelRunner to generate sound”. MCP setup guide.

Model Detayları

Model Detayları

Gemini 3.1 Flash TTS, Google’ın kontrol edilebilir text-to-speech modelidir. Yazılı metni doğal ve ifadeli bir konuşmaya dönüştürür; performansı markup ya da SSML yerine düz bir dille yönlendirmenize izin verir.

`prompt` alanı hem stil yönergesini hem de söylenecek sözleri taşır; biçimi `{style instruction}: {text}` şeklindedir — örneğin `Say the following in a warm, curious way: OK, so... tell me about this AI thing.` Model bu yönergeyi aksan, tempo, ton ve duygusal ifade için izler; böylece aynı cümle hiçbir şeyi yeniden kaydetmeden fısıltı, haber spikeri okuması, heyecanlı bir ara söz ya da sakin bir sesli kitap anlatımı olarak seslendirilebilir. Onu klasik bir konuşma sentezleyicisine tercih etmenin başlıca sebebi budur: okuyuş, sesin sabit bir özelliği değil, bir parametredir.

Otuz hazır ses sunulur, her birinin kendi karakteri vardır — Kore kararlı, Puck neşeli, Aoede rahat, Charon bilgilendirici, Sulafat sıcak, Enceladus fısıltılı, Gacrux olgun, Zubenelgenubi senli benli — ve hepsi `language_code` ile desteklenen bir locale’e eşlenebilir. Gemini 3.1 ayrıca metnin içine yerleştirebileceğiniz `[laughs]` ya da `[sigh]` gibi ifadeli ses etiketleri sunar; bunlar tek başına bir stil cümlesinin verdiğinden daha ince bir anlatım kontrolü sağlar.

Prompt ve metin toplamda kabaca 8,000 byte ile sınırlıdır ve tek bir çağrı yaklaşık 655 saniyeye kadar ses döndürür; bundan uzun girdi kırpılır. Çıktı 24 kHz mono bir WAV dosyasıdır ve faturalandırma üretilen sesin saniyesi üzerindendir; bu yüzden kısa bir replik bir sentin küçük bir kesri kadar tutar.

Seslendirme ve anlatım, sesli kitap ve e-öğrenme, IVR ve asistan anonsları, karakter diyalogları için ve genel bir okuyuş yerine belirli bir okuyuşa ihtiyaç duyduğunuz her yerde kullanın.