minimax/speech-2.8-turbo

ドキュメント
スキーマ

[Core Function] MiniMax Speech 2.8 Turbo is a lower-latency text-to-speech model with the same control surface as Speech 2.8 HD, including paralinguistic tags such as (laughs). [Strengths] Faster and more cost-efficient synthesis while retaining prosody, timbre mix, pronunciation, subtitle, and audio-format controls. [Best For] Highly recommended for: interactive assistants, high-volume TTS batches, cost-sensitive voiceovers, quick narration drafts, and latency-sensitive product prompts. [Limitations] Do NOT use this for streaming or real-time partial audio. Do NOT request hex output or emotion whisper. Do NOT send both voice_id and timbre_weights. Prefer speech-2.8-hd if maximum audio quality is required. Audio URLs expire in about 24 hours. [Routing] Choose speech-2.8-turbo when the user emphasizes speed, cost, or throughput. Choose speech-2.8-hd for premium quality or highly expressive delivery.

$69.0000/M chars
text-to-speech

入力

Text to synthesize (fewer than 10000 characters). Supports pause markers such as <#0.5#>, inline pronunciation in parentheses, and paralinguistic tags such as (laughs) or (sighs).
Required when not using timbre_weights (exactly one of voice_id or timbre_weights). Accepts a system, cloned, or AI-generated voice ID. Mutually exclusive with timbre_weights. Full official system voice catalog: https://platform.minimax.io/docs/faq/system-voice-id . Sample system voices (not exhaustive): English_expressive_narrator, English_radiant_girl, English_magnetic_voiced_man, Chinese (Mandarin)_News_Anchor, Chinese (Mandarin)_Lyrical_Voice, Chinese (Mandarin)_Reliable_Executive.
Custom pronunciation rules in original/replacement form, for example omg/oh my god or Chinese pinyin with tone numbers.
Audio bitrate. Applies to mp3 only.
Number of channels. 1 is mono, 2 is stereo.
Optional emotion control. Leave unset to let the model choose naturally. whisper is not supported on speech-2.8-hd or speech-2.8-turbo.
Force constant bitrate encoding. Optional; typically unused for standard async synthesis.
Output audio format. pcmu_raw and pcmu_wav are G.711 mu-law (8 kHz).
Boost recognition for a language or dialect. Use auto when the language is unknown.
Enable LaTeX formula reading (Chinese). When enabled, language_boost may be forced to Chinese. Wrap formulas with $$.
Voice effect: stronger (negative) to softer (positive).
Voice effect: deepen (negative) to brighten (positive). Distinct from pitch (semitone).
Voice effect: fuller (negative) to crisper (positive).
Pitch adjustment in semitones. Distinct from modify_pitch (voice effect slider).
Output sample rate in Hz.
Optional single sound effect applied to the output.
Speech rate multiplier. Higher is faster.
Whether to generate a subtitle file download link.
Subtitle granularity. word_streaming is not supported.
Enable Chinese/English text normalization for better number reading; may add slight latency.
Required when not using voice_id (exactly one of voice_id or timbre_weights). Mix up to 4 voices by weight. Mutually exclusive with top-level voice_id.
項目 1
voice_id*

Voice ID for one entry in a timbre mix. Distinct from the top-level single-voice voice_id field. May be a system voice from the official catalog (https://platform.minimax.io/docs/faq/system-voice-id) or a cloned/AI voice ID.

weight*

Relative mix weight for this voice. Higher values increase similarity to this voice. At most 4 voices.

Speech volume. Range (0, 10].

結果

結果はまだありません

モデルを実行すると、ここで出力をプレビューできます。

次へ:

README

属性

パラメーター

名前 説明 必須 列挙値
text Text to synthesize (fewer than 10000 characters). Supports pause markers such as <#0.5#>, inline pronunciation in parentheses, and paralinguistic tags such as (laughs) or (sighs). string はい -
voice_id Required when not using timbre_weights (exactly one of voice_id or timbre_weights). Accepts a system, cloned, or AI-generated voice ID. Mutually exclusive with timbre_weights. Full official system voice catalog: https://platform.minimax.io/docs/faq/system-voice-id . Sample system voices (not exhaustive): English_expressive_narrator, English_radiant_girl, English_magnetic_voiced_man, Chinese (Mandarin)_News_Anchor, Chinese (Mandarin)_Lyrical_Voice, Chinese (Mandarin)_Reliable_Executive. string いいえ -
pronunciation_tone Custom pronunciation rules in original/replacement form, for example omg/oh my god or Chinese pinyin with tone numbers. string[] いいえ -
bitrate Audio bitrate. Applies to mp3 only. integer いいえ 32000, 64000, 128000, 256000
channel Number of channels. 1 is mono, 2 is stereo. integer いいえ 1, 2
emotion Optional emotion control. Leave unset to let the model choose naturally. whisper is not supported on speech-2.8-hd or speech-2.8-turbo. string いいえ happy, sad, angry, fearful, disgusted, surprised, calm, fluent
force_cbr Force constant bitrate encoding. Optional; typically unused for standard async synthesis. boolean いいえ true, false
format Output audio format. pcmu_raw and pcmu_wav are G.711 mu-law (8 kHz). string いいえ mp3, pcm, flac, wav, pcmu_raw, pcmu_wav, opus
language_boost Boost recognition for a language or dialect. Use auto when the language is unknown. string いいえ Chinese, Chinese,Yue, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Persian, Slovak, Swedish, Croatian, Filipino, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Tamil, Afrikaans, auto
latex_read Enable LaTeX formula reading (Chinese). When enabled, language_boost may be forced to Chinese. Wrap formulas with $$. boolean いいえ true, false
modify_intensity Voice effect: stronger (negative) to softer (positive). integer いいえ -
modify_pitch Voice effect: deepen (negative) to brighten (positive). Distinct from pitch (semitone). integer いいえ -
modify_timbre Voice effect: fuller (negative) to crisper (positive). integer いいえ -
pitch Pitch adjustment in semitones. Distinct from modify_pitch (voice effect slider). integer いいえ -
sample_rate Output sample rate in Hz. integer いいえ 8000, 16000, 22050, 24000, 32000, 44100
sound_effects Optional single sound effect applied to the output. string いいえ spacious_echo, auditorium_echo, lofi_telephone, robotic
speed Speech rate multiplier. Higher is faster. number いいえ -
subtitle_enable Whether to generate a subtitle file download link. boolean いいえ true, false
subtitle_type Subtitle granularity. word_streaming is not supported. string いいえ sentence, word
text_normalization Enable Chinese/English text normalization for better number reading; may add slight latency. boolean いいえ true, false
timbre_weights Required when not using voice_id (exactly one of voice_id or timbre_weights). Mix up to 4 voices by weight. Mutually exclusive with top-level voice_id. object[] いいえ -
vol Speech volume. Range (0, 10]. number いいえ -

料金

単位: $/M chars

料金
$69.0000/M chars