alibaba/qwen-audio-3.0-tts-plus

qwen-audio-3.0-tts-plus
ドキュメント
スキーマ

[Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanlingxin and longanlufeng (see the Qwen-Audio-TTS voice list). [Best For] Premium narration, brand voiceovers, multilingual product audio, and quality-sensitive batch TTS when Qwen-Audio voices are preferred. [Limitations] Do NOT mix Flash-only voices (e.g. longanhuan_v3.6) with Plus. [Routing] Choose Plus when speech quality is the priority. Choose qwen-audio-3.0-tts-flash when lower latency matters more.

$12.6500/M chars
text-to-speech

入力

Text to synthesize. Required. Max 20,000 Unicode characters. Supports plain text and SSML (set enable_ssml to true) per Alibaba documentation.
Voice ID for qwen-audio-3.0-tts-plus only. Use a Plus system voice (e.g. longanlingxin, longanlufeng) or a Plus base/cloned voice ID from the official Qwen-Audio-TTS voice list: https://help.aliyun.com/zh/model-studio/qwen-audio-tts-voice-list . Do not mix Flash system voices with Plus.
AIGC PropagateID. Only effective when enable_aigc_tag is true.
AIGC ContentPropagator. Only effective when enable_aigc_tag is true.
Optional speaking-style instruction. Enforced weighted length ≤100 (CJK ideographs / Han count as 2; other characters count as 1).
Audio bit rate in kbps. Optional. Range 6 to 510. Only supported when format is opus; do not use for mp3, pcm, or wav.
Embed AIGC invisible watermark into wav/mp3/opus output. Supported on Qwen-Audio 3.0 TTS Plus and Flash.
Whether to parse text as SSML. Default false. When true, text must follow Alibaba SpeechSynthesizer SSML rules.
Audio encoding format. Default mp3.
Target language hint for pronunciation (numbers, symbols, minor languages).
Pitch multiplier. Default 1.0. Range 0.5 (lower) to 2.0 (higher).
Speech rate multiplier. Default 1.0. Range 0.5 (slow) to 2.0 (fast).
Audio sample rate in Hz. Default 22050.
Random seed for reproducible synthesis when text, voice, and other parameters are identical. Default 0. Range 0 to 65535.
Output volume. Default 50. Range 0 (silent) to 100 (maximum).

結果

結果はまだありません

モデルを実行すると、ここで出力をプレビューできます。

次へ:

README

属性

シリーズ

パラメーター

名前 説明 必須 列挙値
text Text to synthesize. Required. Max 20,000 Unicode characters. Supports plain text and SSML (set enable_ssml to true) per Alibaba documentation. string はい -
voice Voice ID for qwen-audio-3.0-tts-plus only. Use a Plus system voice (e.g. longanlingxin, longanlufeng) or a Plus base/cloned voice ID from the official Qwen-Audio-TTS voice list: https://help.aliyun.com/zh/model-studio/qwen-audio-tts-voice-list . Do not mix Flash system voices with Plus. string はい -
aigc_propagate_id AIGC PropagateID. Only effective when enable_aigc_tag is true. string いいえ -
aigc_propagator AIGC ContentPropagator. Only effective when enable_aigc_tag is true. string いいえ -
instruction Optional speaking-style instruction. Enforced weighted length ≤100 (CJK ideographs / Han count as 2; other characters count as 1). string いいえ -
bit_rate Audio bit rate in kbps. Optional. Range 6 to 510. Only supported when format is opus; do not use for mp3, pcm, or wav. integer いいえ -
enable_aigc_tag Embed AIGC invisible watermark into wav/mp3/opus output. Supported on Qwen-Audio 3.0 TTS Plus and Flash. boolean いいえ true, false
enable_ssml Whether to parse text as SSML. Default false. When true, text must follow Alibaba SpeechSynthesizer SSML rules. boolean いいえ true, false
format Audio encoding format. Default mp3. string いいえ mp3, pcm, wav, opus
language_hint Target language hint for pronunciation (numbers, symbols, minor languages). string いいえ zh, en, fr, de, ja, ko, ru, pt, th, id, vi, es, it, ms, fil, ar
pitch Pitch multiplier. Default 1.0. Range 0.5 (lower) to 2.0 (higher). number いいえ -
rate Speech rate multiplier. Default 1.0. Range 0.5 (slow) to 2.0 (fast). number いいえ -
sample_rate Audio sample rate in Hz. Default 22050. integer いいえ 8000, 16000, 22050, 24000, 44100, 48000
seed Random seed for reproducible synthesis when text, voice, and other parameters are identical. Default 0. Range 0 to 65535. integer いいえ -
volume Output volume. Default 50. Range 0 (silent) to 100 (maximum). integer いいえ -

料金

単位: $/M chars

料金
$12.6500/M chars

関連モデル

  • alibaba/qwen-audio-3.0-tts-flash: [Core Function] Qwen-Audio 3.0 TTS Flash is Alibaba’s low-latency Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Fast synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanhuan_v3.6, longjielidou_v3.6, loongeva_v3.6, and loongjohn (see the Qwen-Audio-TTS voice list). [Best For] Voice assistants, interactive prompts, short announcements, multilingual product flows, and latency-sensitive batch TTS using Qwen-Audio voices. [Limitations] Do NOT mix Plus-only voices (e.g. longanlingxin) with Flash. Use Plus when maximum narration quality matters more than turnaround time. [Routing] Choose Flash when speed matters most. Choose qwen-audio-3.0-tts-plus for premium narration quality.

関連リソース