Series

Qwen Audio AI Model Family

2 modelsUpdated Jul 2026
Qwen Audio AI Model Family

About Qwen Audio Models

Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.

All Qwen Audio Models

alibaba/qwen-audio-3.0-tts-flash

alibaba/qwen-audio-3.0-tts-flash

text-to-speech

[Core Function] Qwen-Audio 3.0 TTS Flash is Alibaba's low-latency Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Fast synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanhuan_v3.6, longjielidou_v3.6, loongeva_v3.6, and loongjohn (see the Qwen-Audio-TTS voice list). [Best For] Voice assistants, interactive prompts, short announcements, multilingual product flows, and latency-sensitive batch TTS using Qwen-Audio voices. [Limitations] Does not expose hot_fix or Markdown filtering. Do NOT mix Plus-only voices (e.g. longanlingxin) with Flash. Use Plus when maximum narration quality matters more than turnaround time. [Routing] Choose Flash when speed matters most. Choose qwen-audio-3.0-tts-plus for premium narration quality.

alibaba/qwen-audio-3.0-tts-plus

alibaba/qwen-audio-3.0-tts-plus

text-to-speech

[Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanlingxin and longanlufeng (see the Qwen-Audio-TTS voice list). [Best For] Premium narration, brand voiceovers, multilingual product audio, and quality-sensitive batch TTS when Qwen-Audio voices are preferred. [Limitations] Does not expose hot_fix or Markdown filtering. Do NOT use for real-time streaming or word-level timestamps. Do NOT mix Flash-only voices (e.g. longanhuan_v3.6) with Plus. [Routing] Choose Plus when speech quality is the priority. Choose qwen-audio-3.0-tts-flash when lower latency matters more.

Can't find the model you need? Let us know.

Explore More

Text to Speech
Category9 models
Voice Cloning
Voice Cloning
Collection2 models
Fun
Fun
Series2 models
CosyVoice
CosyVoice
Series4 models
Grok Voice
Grok Voice
Series2 models
Speech to Text
Category5 models