Series

CosyVoice AI Model Family

4 modelsUpdated Jul 2026
CosyVoice AI Model Family

About CosyVoice Models

CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.

All CosyVoice Models

alibaba/cosyvoice-design

alibaba/cosyvoice-design

text-to-speech

[Core Function] CosyVoice Design creates a temporary voice from a natural-language voice_prompt and synthesizes speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint (zh/en) applies to both design and synthesis; same prosody and format controls as CosyVoice TTS. [Best For] One-off designed speech without storing enrolled voices. [Limitations] Do NOT use this to obtain a reusable voice library entry; the designed voice is temporary and is not returned. model must be cosyvoice-v3.5-plus or cosyvoice-v3.5-flash; voice_prompt and text are required.

alibaba/cosyvoice-clone

alibaba/cosyvoice-clone

speech-to-speech

[Core Function] CosyVoice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint applies to both cloning and synthesis; SSML, hot_fix, and prosody controls. [Best For] One-off cloned narration and demos where a lasting voice library is not needed. [Limitations] Do NOT use this to obtain a reusable voice library entry; the cloned voice is temporary and is not returned. Reference URL must be publicly accessible. model must be cosyvoice-v3.5-plus or cosyvoice-v3.5-flash. [Routing] Choose cosyvoice-v3.5-plus for higher speech quality; cosyvoice-v3.5-flash for lower latency.

alibaba/cosyvoice-v3-flash

alibaba/cosyvoice-v3-flash

text-to-speech

[Core Function] CosyVoice v3 Flash is Alibaba's low-latency text-to-speech model. [Strengths] Rich system voice catalog, fixed-format instruction on Instruct-capable system voices, SSML, hot_fix, AIGC watermark, Markdown filter (cloned voices only), and multiple audio formats with faster turnaround than Plus. [Best For] Voice assistants, interactive prompts, IVR, short announcements, dialect system voices, and latency-sensitive batch TTS. [Limitations] Do NOT use when maximum speech quality or long-form audiobook fidelity is the priority (use Plus). System-voice instruction must follow CosyVoice voice-list fixed Chinese formats. [Routing] Choose Flash when speed or a richer system-voice catalog matters most. Choose Plus for premium narration quality.

alibaba/cosyvoice-v3-plus

alibaba/cosyvoice-v3-plus

text-to-speech

[Core Function] CosyVoice v3 Plus is Alibaba's high-quality text-to-speech model. [Strengths] System voices (e.g. longanyang, longanhuan), SSML and LaTeX input, hot_fix pronunciation correction, AIGC watermark, and output in mp3, pcm, wav, or opus. System-voice instruction must use the fixed Chinese formats in the CosyVoice voice list. [Best For] Brand voiceovers, audiobooks, high-quality narration, marketing clips, and scenarios where speech quality matters more than minimum latency. [Limitations] Do NOT use for real-time streaming or word-level timestamps. Fewer system voices than Flash. [Routing] Choose Plus when quality or narration fidelity matters most. Choose cosyvoice-v3-flash for lower latency or a richer system-voice catalog.

Can't find the model you need? Let us know.

Explore More

Qwen Audio
Qwen Audio
Series2 models
Text to Speech
Category9 models
Voice Cloning
Voice Cloning
Collection2 models
Fun
Fun
Series2 models
Grok Voice
Grok Voice
Series2 models
Speech to Text
Category5 models