microsoft/mai-transcribe-1.5

ドキュメント
スキーマ

[Core Function] MAI-Transcribe 1.5 transcribes a single public audio URL into text via an async task. [Strengths] Multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps. [Best For] Meeting notes, captions, and batch audio-to-text. [Limitations] Do NOT use file upload; URL-only input. Do NOT expect speaker diarization. Supported audio: WAV, MP3, or FLAC up to 300 MB. [Routing] Use this model for Microsoft MAI speech recognition quality with URL-based audio.

$0.0001/sec
speech-to-text

入力

Public http(s) URL of the audio file. Supported formats: WAV, MP3, FLAC (up to 300 MB).
ヒント:ファイルをドラッグ&ドロップするか、クリップボード(Ctrl/Cmd+V)またはURLから追加できます。
Domain phrases for entity biasing (product names, proper nouns).
Optional language codes to force recognition. Omit for multi-lingual auto mode.
locales

結果

結果はまだありません

モデルを実行すると、ここで出力をプレビューできます。

次へ:

README

属性

パラメーター

名前 説明 必須 列挙値
url Public http(s) URL of the audio file. Supported formats: WAV, MP3, FLAC (up to 300 MB). string はい -
phrase_list Domain phrases for entity biasing (product names, proper nouns). string[] いいえ -
locales Optional language codes to force recognition. Omit for multi-lingual auto mode. string[] いいえ ar, as, bg, bn, ca, cs, da, de, el, en, es, et, fi, fr, gu, hi, hu, id, it, ja, kn, ko, lt, ml, mr, nb, nl, or, pa, pl, pt, ro, ru, sk, sl, sv, ta, te, th, tr, uk, vi, zh

料金

単位: $/sec

料金
$0.0001/sec