分类

2026 年最佳语音转语音 AI 模型

2 个模型更新于 Sep 2026

关于 语音转语音 模型

在 Modellix 探索 2 个生产可用的语音转语音 AI 模型,对比模型能力、在线试用,并通过统一 API 快速完成集成。

全部 语音转语音 模型

minimax/minimax-voice-clone

minimax/minimax-voice-clone

speech-to-speech

[Core Function] MiniMax Voice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; optional language_boost for clone and synthesis; prosody controls (speed, volume, pitch, emotion); pronunciation overrides; voice effects; flexible audio formats. [Best For] One-off cloned narration, personalized prompts, and demos where a lasting voice library is not needed. [Limitations] Do NOT use this to obtain a reusable voice library entry; the cloned voice is temporary and is not returned. audio_url must be publicly reachable (mp3/m4a/wav, 10s–5min, ≤20 MB). [Routing] Choose speech-2.8-hd for higher quality; speech-2.8-turbo for lower latency or cost. Paralinguistic tags such as (laughs) in text are supported on both speech-2.8-hd and speech-2.8-turbo.

alibaba/cosyvoice-clone

alibaba/cosyvoice-clone

speech-to-speech

[Core Function] CosyVoice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint applies to both cloning and synthesis; SSML, hot_fix, and prosody controls. [Best For] One-off cloned narration and demos where a lasting voice library is not needed. [Limitations] Do NOT use this to obtain a reusable voice library entry; the cloned voice is temporary and is not returned. Reference URL must be publicly accessible. model must be cosyvoice-v3.5-plus or cosyvoice-v3.5-flash. [Routing] Choose cosyvoice-v3.5-plus for higher speech quality; cosyvoice-v3.5-flash for lower latency.

没有找到需要的模型? 告诉我们。

探索更多

语音转文本
分类6 个模型
Qwen Audio
Qwen Audio
系列2 个模型
文生语音
分类9 个模型
Voice Cloning
Voice Cloning
集合2 个模型
Fun
Fun
系列2 个模型
CosyVoice
CosyVoice
系列4 个模型