google/gemini-3.1-flash-tts

ドキュメント
スキーマ

[Core Function] Gemini 3.1 Flash TTS is Google's low-latency, controllable text-to-speech model. [Strengths] Single-speaker and two-speaker dialogue, 30 prebuilt voices, 70+ languages via language_code, and expressive delivery through style prompts plus inline audio tags such as [whispers], [slow], [fast], and [laughs]. Output is WAV (24 kHz mono PCM). [Best For] Voiceovers, virtual presenters, audiobook narration, multilingual speech, podcast-style scripts, and two-person dialogue. [Limitations] Do NOT use for lip-syncing existing video, music or sound-effect generation, image or audio inputs, MP3/OGG export, or API-level speed, volume, encoding, or sample-rate controls. Combined prompt and text must stay within 8,000 bytes. [Routing] Route here for controllable Gemini TTS from text only; for video with native audio use Veo; for talking-head lip sync from a portrait plus audio use Kling Avatar or SkyReels avatar models.

$16.1000/M chars
text-to-speech

入力

Optional director or style instruction (tone, pace, emotion, scene). Max 4,000 bytes. Merged with text as "{prompt}: {text}". Combined prompt + ": " + text must not exceed 8,000 bytes. If you omit prompt, you may embed direction in text using Google's single-field pattern (e.g. "Say cheerfully: ...") or audio tags.
Script to be spoken. Required. Max 8,000 bytes when used alone; if prompt is set, combined "{prompt}: {text}" must not exceed 8,000 bytes. When prompt is set, put only the spoken script here (do not repeat style instructions such as "Say cheerfully:" in text). May include inline audio tags such as [whispers], [slow], [fast], [laughs], and [short pause] to steer delivery.
Speaker configuration. Required. Length 1 (single voice) or 2 (dialogue). Voice is always set here; there is no top-level voice field.
項目 1
speaker*

Speaker label. For dialogue, must match names used in text (e.g. "Joe: ..."). Alphanumeric only, 1-64 characters.

voice*

Prebuilt Gemini TTS voice name.

Optional BCP-47 language and accent code. Examples: en-us, en-in, cmn-cn, ja-jp, ko-kr, fr-fr.

結果

結果はまだありません

モデルを実行すると、ここで出力をプレビューできます。

次へ:

README

属性

パラメーター

名前 説明 必須 列挙値
prompt Optional director or style instruction (tone, pace, emotion, scene). Max 4,000 bytes. Merged with text as “{prompt}: {text}”. Combined prompt + ": " + text must not exceed 8,000 bytes. If you omit prompt, you may embed direction in text using Google’s single-field pattern (e.g. “Say cheerfully: …”) or audio tags. string いいえ -
text Script to be spoken. Required. Max 8,000 bytes when used alone; if prompt is set, combined “{prompt}: {text}” must not exceed 8,000 bytes. When prompt is set, put only the spoken script here (do not repeat style instructions such as “Say cheerfully:” in text). May include inline audio tags such as [whispers], [slow], [fast], [laughs], and [short pause] to steer delivery. string はい -
speakers Speaker configuration. Required. Length 1 (single voice) or 2 (dialogue). Voice is always set here; there is no top-level voice field. object[] はい -
language_code Optional BCP-47 language and accent code. Examples: en-us, en-in, cmn-cn, ja-jp, ko-kr, fr-fr. string いいえ -

料金

単位: $/M chars

料金
$16.1000/M chars