microsoft/mai-transcribe-1.5

文档
Schema

[Core Function] MAI-Transcribe 1.5 transcribes a single public audio URL into text via an async task. [Strengths] Multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps. [Best For] Meeting notes, captions, and batch audio-to-text. [Limitations] Do NOT use file upload; URL-only input. Do NOT expect speaker diarization. Supported audio: WAV, MP3, or FLAC up to 300 MB. [Routing] Use this model for Microsoft MAI speech recognition quality with URL-based audio.

$0.0001/sec
speech-to-text

输入

Public http(s) URL of the audio file. Supported formats: WAV, MP3, FLAC (up to 300 MB).
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Domain phrases for entity biasing (product names, proper nouns).
Optional language codes to force recognition. Omit for multi-lingual auto mode.
locales

结果

暂无结果

运行模型后,结果将在这里显示。

Next:

README

属性

参数

参数名 描述 类型 必填 枚举值
url Public http(s) URL of the audio file. Supported formats: WAV, MP3, FLAC (up to 300 MB). string -
phrase_list Domain phrases for entity biasing (product names, proper nouns). string[] -
locales Optional language codes to force recognition. Omit for multi-lingual auto mode. string[] ar, as, bg, bn, ca, cs, da, de, el, en, es, et, fi, fr, gu, hi, hu, id, it, ja, kn, ko, lt, ml, mr, nb, nl, or, pa, pl, pt, ro, ru, sk, sl, sv, ta, te, th, tr, uk, vi, zh

价格

单位: $/sec

价格
$0.0001/sec