分类

2026 年最佳语音转文本 AI 模型

5 个模型更新于 Jul 2026

关于 语音转文本 模型

在 Modellix 探索 5 个生产可用的语音转文本 AI 模型,对比模型能力、在线试用,并通过统一 API 快速完成集成。

全部 语音转文本 模型

xai/grok-voice-asr

xai/grok-voice-asr

speech-to-text

[Core Function] Grok Voice ASR transcribes a single public audio URL into text via an async task. [Strengths] Word-level timestamps, optional speaker diarization, multichannel transcription, Inverse Text Normalization (format + language), keyterm biasing, and filler-word control. [Best For] Meeting notes, call-center recordings, captions, and batch audio-to-text pipelines. [Limitations] Do NOT use file upload; URL-only input. Do NOT use for live or real-time streaming transcription. Audio must be publicly reachable (max 500 MB). [Routing] Use this model for xAI Grok Voice ASR quality with URL-based audio.

openai/whisper-1

openai/whisper-1

speech-to-text

[Core Function] OpenAI Whisper transcribes a single public audio URL into text via an async task. [Strengths] Multiple output formats (verbose_json with word/segment timestamps, plain text, SRT, VTT), optional language and prompt biasing. [Best For] Meeting notes, podcasts, captions, and batch audio-to-text. [Limitations] Do NOT use file upload; URL-only input (audio up to 25 MB). Prefer verbose_json when you need duration or timestamps. [Routing] Use whisper-1 for OpenAI Whisper quality with URL-based audio.

microsoft/mai-transcribe-1.5

microsoft/mai-transcribe-1.5

speech-to-text

[Core Function] MAI-Transcribe 1.5 transcribes a single public audio URL into text via an async task. [Strengths] Multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps. [Best For] Meeting notes, captions, and batch audio-to-text. [Limitations] Do NOT use file upload; URL-only input. Do NOT expect speaker diarization. Supported audio: WAV, MP3, or FLAC up to 300 MB. [Routing] Use this model for Microsoft MAI speech recognition quality with URL-based audio.

alibaba/fun-asr-mtl

alibaba/fun-asr-mtl

speech-to-text

[Core Function] Fun-ASR MTL is the multi-language variant for async recorded speech recognition. [Strengths] Same parameters as fun-asr with multi-language tuning. [Best For] Mixed-language or international audio archives. [Limitations] Do NOT send more than one file per request. [Routing] Choose fun-asr for general use; fun-asr-mtl when the source audio is explicitly multi-language.

alibaba/fun-asr

alibaba/fun-asr

speech-to-text

[Core Function] Fun-ASR transcribes a single public audio file asynchronously. [Strengths] Hot-word vocabulary, optional speaker diarization, channel selection, and language hints. [Best For] Batch transcription of recordings up to 12 hours. [Limitations] Do NOT send more than one file per request. [Routing] Use fun-asr-mtl when multi-language optimization is preferred.

没有找到需要的模型? 告诉我们。

探索更多