microsoft/mai-transcribe-1.5

Docs
Schema

[Core Function] MAI-Transcribe 1.5 transcribes a single public audio URL into text via an async task. [Strengths] Multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps. [Best For] Meeting notes, captions, and batch audio-to-text. [Limitations] Do NOT use file upload; URL-only input. Do NOT expect speaker diarization. Supported audio: WAV, MP3, or FLAC up to 300 MB. [Routing] Use this model for Microsoft MAI speech recognition quality with URL-based audio.

$0.0001/sec
speech-to-text

Input

Public http(s) URL of the audio file. Supported formats: WAV, MP3, FLAC (up to 300 MB).
Hint: Drag and drop files, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Domain phrases for entity biasing (product names, proper nouns).
Optional language codes to force recognition. Omit for multi-lingual auto mode.
locales

Result

No results yet

Run the model to preview the output here.

Next:

README

Attributes

Parameters

Name Description Type Required Enums
url Public http(s) URL of the audio file. Supported formats: WAV, MP3, FLAC (up to 300 MB). string Yes -
phrase_list Domain phrases for entity biasing (product names, proper nouns). string[] No -
locales Optional language codes to force recognition. Omit for multi-lingual auto mode. string[] No ar, as, bg, bn, ca, cs, da, de, el, en, es, et, fi, fr, gu, hi, hu, id, it, ja, kn, ko, lt, ml, mr, nb, nl, or, pa, pl, pt, ro, ru, sk, sl, sv, ta, te, th, tr, uk, vi, zh

Pricing

Unit: $/sec

Pricing
$0.0001/sec