minimax/speech-2.8-turbo

文档
Schema

[Core Function] MiniMax Speech 2.8 Turbo is a lower-latency text-to-speech model with the same control surface as Speech 2.8 HD, including paralinguistic tags such as (laughs). [Strengths] Faster and more cost-efficient synthesis while retaining prosody, timbre mix, pronunciation, subtitle, and audio-format controls. [Best For] Highly recommended for: interactive assistants, high-volume TTS batches, cost-sensitive voiceovers, quick narration drafts, and latency-sensitive product prompts. [Limitations] Do NOT use this for streaming or real-time partial audio. Do NOT request hex output or emotion whisper. Do NOT send both voice_id and timbre_weights. Prefer speech-2.8-hd if maximum audio quality is required. Audio URLs expire in about 24 hours. [Routing] Choose speech-2.8-turbo when the user emphasizes speed, cost, or throughput. Choose speech-2.8-hd for premium quality or highly expressive delivery.

$69.0000/M chars
text-to-speech

输入

Text to synthesize (fewer than 10000 characters). Supports pause markers such as <#0.5#>, inline pronunciation in parentheses, and paralinguistic tags such as (laughs) or (sighs).
Required when not using timbre_weights (exactly one of voice_id or timbre_weights). Accepts a system, cloned, or AI-generated voice ID. Mutually exclusive with timbre_weights. Full official system voice catalog: https://platform.minimax.io/docs/faq/system-voice-id . Sample system voices (not exhaustive): English_expressive_narrator, English_radiant_girl, English_magnetic_voiced_man, Chinese (Mandarin)_News_Anchor, Chinese (Mandarin)_Lyrical_Voice, Chinese (Mandarin)_Reliable_Executive.
Custom pronunciation rules in original/replacement form, for example omg/oh my god or Chinese pinyin with tone numbers.
Audio bitrate. Applies to mp3 only.
128000
Number of channels. 1 is mono, 2 is stereo.
1
Optional emotion control. Leave unset to let the model choose naturally. whisper is not supported on speech-2.8-hd or speech-2.8-turbo.
emotion
Force constant bitrate encoding. Optional; typically unused for standard async synthesis.
Output audio format. pcmu_raw and pcmu_wav are G.711 mu-law (8 kHz).
mp3
Boost recognition for a language or dialect. Use auto when the language is unknown.
language_boost
Enable LaTeX formula reading (Chinese). When enabled, language_boost may be forced to Chinese. Wrap formulas with $$.
Voice effect: stronger (negative) to softer (positive).
Voice effect: deepen (negative) to brighten (positive). Distinct from pitch (semitone).
Voice effect: fuller (negative) to crisper (positive).
Pitch adjustment in semitones. Distinct from modify_pitch (voice effect slider).
Output sample rate in Hz.
32000
Optional single sound effect applied to the output.
sound_effects
Speech rate multiplier. Higher is faster.
Whether to generate a subtitle file download link.
Subtitle granularity. word_streaming is not supported.
sentence
Enable Chinese/English text normalization for better number reading; may add slight latency.
Required when not using voice_id (exactly one of voice_id or timbre_weights). Mix up to 4 voices by weight. Mutually exclusive with top-level voice_id.
第 1 项
voice_id*

Voice ID for one entry in a timbre mix. Distinct from the top-level single-voice voice_id field. May be a system voice from the official catalog (https://platform.minimax.io/docs/faq/system-voice-id) or a cloned/AI voice ID.

weight*

Relative mix weight for this voice. Higher values increase similarity to this voice. At most 4 voices.

Speech volume. Range (0, 10].

结果

暂无结果

运行模型后,结果将在这里显示。

Next:

README

属性

参数

参数名 描述 类型 必填 枚举值
text Text to synthesize (fewer than 10000 characters). Supports pause markers such as <#0.5#>, inline pronunciation in parentheses, and paralinguistic tags such as (laughs) or (sighs). string -
voice_id Required when not using timbre_weights (exactly one of voice_id or timbre_weights). Accepts a system, cloned, or AI-generated voice ID. Mutually exclusive with timbre_weights. Full official system voice catalog: https://platform.minimax.io/docs/faq/system-voice-id . Sample system voices (not exhaustive): English_expressive_narrator, English_radiant_girl, English_magnetic_voiced_man, Chinese (Mandarin)_News_Anchor, Chinese (Mandarin)_Lyrical_Voice, Chinese (Mandarin)_Reliable_Executive. string -
pronunciation_tone Custom pronunciation rules in original/replacement form, for example omg/oh my god or Chinese pinyin with tone numbers. string[] -
bitrate Audio bitrate. Applies to mp3 only. integer 32000, 64000, 128000, 256000
channel Number of channels. 1 is mono, 2 is stereo. integer 1, 2
emotion Optional emotion control. Leave unset to let the model choose naturally. whisper is not supported on speech-2.8-hd or speech-2.8-turbo. string happy, sad, angry, fearful, disgusted, surprised, calm, fluent
force_cbr Force constant bitrate encoding. Optional; typically unused for standard async synthesis. boolean true, false
format Output audio format. pcmu_raw and pcmu_wav are G.711 mu-law (8 kHz). string mp3, pcm, flac, wav, pcmu_raw, pcmu_wav, opus
language_boost Boost recognition for a language or dialect. Use auto when the language is unknown. string Chinese, Chinese,Yue, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Persian, Slovak, Swedish, Croatian, Filipino, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Tamil, Afrikaans, auto
latex_read Enable LaTeX formula reading (Chinese). When enabled, language_boost may be forced to Chinese. Wrap formulas with $$. boolean true, false
modify_intensity Voice effect: stronger (negative) to softer (positive). integer -
modify_pitch Voice effect: deepen (negative) to brighten (positive). Distinct from pitch (semitone). integer -
modify_timbre Voice effect: fuller (negative) to crisper (positive). integer -
pitch Pitch adjustment in semitones. Distinct from modify_pitch (voice effect slider). integer -
sample_rate Output sample rate in Hz. integer 8000, 16000, 22050, 24000, 32000, 44100
sound_effects Optional single sound effect applied to the output. string spacious_echo, auditorium_echo, lofi_telephone, robotic
speed Speech rate multiplier. Higher is faster. number -
subtitle_enable Whether to generate a subtitle file download link. boolean true, false
subtitle_type Subtitle granularity. word_streaming is not supported. string sentence, word
text_normalization Enable Chinese/English text normalization for better number reading; may add slight latency. boolean true, false
timbre_weights Required when not using voice_id (exactly one of voice_id or timbre_weights). Mix up to 4 voices by weight. Mutually exclusive with top-level voice_id. object[] -
vol Speech volume. Range (0, 10]. number -

价格

单位: $/M chars

价格
$69.0000/M chars