Series

Grok Voice AI Model Family

2 modelsUpdated Jul 2026
Grok Voice AI Model Family

About Grok Voice Models

Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.

All Grok Voice Models

xai/grok-voice-tts

xai/grok-voice-tts

text-to-speech

[Core Function] Grok Voice TTS converts text into natural spoken audio with expressive voices and optional speech tags embedded in the text. [Strengths] Supports 20+ languages (plus auto-detect), 26 built-in voices, multiple codecs (mp3/wav/pcm/mulaw/alaw), and custom voice IDs. [Best For] Product voiceovers, IVR/telephony (mulaw/alaw), and multilingual narration. [Limitations] Do NOT exceed 15,000 characters per request. Audio is returned on the completed task after polling. [Routing] Use this model for xAI Grok Voice quality or custom cloned voices. [Built-in voices] Original: eve (default), ara, leo, rex, sal. Flagship: altair, atlas, carina, castor, celeste, cosmo, helios, helix, iris, kepler, lumen, luna, lux, naksh, orion, perseus, rigel, sirius, ursa, zagan, zenith — case-insensitive; custom voice IDs are also accepted as voice_id.

xai/grok-voice-asr

xai/grok-voice-asr

speech-to-text

[Core Function] Grok Voice ASR transcribes a single public audio URL into text via an async task. [Strengths] Word-level timestamps, optional speaker diarization, multichannel transcription, Inverse Text Normalization (format + language), keyterm biasing, and filler-word control. [Best For] Meeting notes, call-center recordings, captions, and batch audio-to-text pipelines. [Limitations] Do NOT use file upload; URL-only input. Do NOT use for live or real-time streaming transcription. Audio must be publicly reachable (max 500 MB). [Routing] Use this model for xAI Grok Voice ASR quality with URL-based audio.

Can't find the model you need? Let us know.

Explore More

Qwen Audio
Qwen Audio
Series2 models
Text to Speech
Category9 models
Voice Cloning
Voice Cloning
Collection2 models
Fun
Fun
Series2 models
CosyVoice
CosyVoice
Series4 models
Speech to Text
Category5 models