Provider

MiniMax AI Models & API in 2026

13 modelsUpdated Aug 2026
MiniMax AI Models & API in 2026

About MiniMax Models

MiniMax is a leading provider of advanced AI media models, featuring Hailuo for high-fidelity multimodal video generation and editing.

All MiniMax Models

Image to Video

minimax/minimax-h3-i2v

minimax/minimax-h3-i2v

image-to-video

[Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions or when the primary input is a reference video. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 I2V when the user provides reference stills. For start or end frames, use MiniMax H3 FL2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

minimax/minimax-h3-fl2v

minimax/minimax-h3-fl2v

image-to-video

[Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 I2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

minimax/hailuo-02-fl2v

minimax/hailuo-02-fl2v

image-to-video

[Core Function] Hailuo 02 FL2V is a First-Last frame transition video model. [Strengths] It excels at generating a logical, physically accurate video transition that bridges a provided starting frame and an ending frame. [Best For] Highly recommended for: visual morphing, before-and-after transitions, and precise narrative storyboard completion. [Limitations] Do NOT use this model if you only have one image (use standard I2V instead). Note that Hailuo 2.3 does not support FL2V, so this is the primary transition model. [Routing] Use this model exclusively when the user provides BOTH a first frame and a last frame for a transition.

minimax/hailuo-02-i2v

minimax/hailuo-02-i2v

image-to-video

[Core Function] Hailuo 02 I2V is an image-to-video model optimized for physical realism and sustained high resolution. [Strengths] It excels at animating broad scenes, maintaining complex physics, and supporting native 1080p generation for up to 10 seconds. [Best For] Highly recommended for: animating product photography, bringing landscape/nature photos to life, and generating physically accurate motion. [Limitations] Do NOT use this model for animating complex human facial micro-expressions or highly stylized anime art, where Hailuo 2.3 is superior. [Routing] Use this model when the user needs to animate a landscape/product, or explicitly requires 10 seconds of 1080p video. Otherwise, default to Hailuo 2.3 I2V.

minimax/hailuo-2.3-fast-i2v

minimax/hailuo-2.3-fast-i2v

image-to-video

[Core Function] Hailuo 2.3 Fast I2V is a high-speed, cost-effective image-to-video generation model. [Strengths] It excels at generating videos from images much faster and at roughly 50% lower cost than the standard 2.3 model, while still maintaining the 2.3 architecture's strength in human motion. [Best For] Highly recommended for: rapid prototyping, batch social media creation, and cost-sensitive video generation pipelines. [Limitations] Do NOT use this model for text-to-video (it only accepts image inputs). Do NOT use when absolute maximum visual fidelity is the primary requirement. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or 'cost-effective' image-to-video generation. For maximum quality, use the standard Hailuo 2.3 I2V.

minimax/hailuo-2.3-i2v

minimax/hailuo-2.3-i2v

image-to-video

[Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.

Text to Video

minimax/minimax-h3-t2v

minimax/minimax-h3-t2v

text-to-video

[Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.

minimax/hailuo-02-t2v

minimax/hailuo-02-t2v

text-to-video

[Core Function] Hailuo 02 T2V is a text-to-video generation model optimized for physical realism. [Strengths] It excels at complex physics simulation, fluid dynamics, broad cinematic scenes, and natively rendering 1080p video up to 10 seconds without downscaling. [Best For] Highly recommended for: product commercials, high-speed sports action, nature documentaries, and sweeping landscapes. [Limitations] Do NOT use this model for highly stylized anime/art or nuanced human micro-expressions, where Hailuo 2.3 performs better. [Routing] Route to this model when the user requests '1080p for 10 seconds', complex physical action (like splashing water or crashes), or broad landscapes. For human characters and stylization, use Hailuo 2.3 T2V.

minimax/hailuo-2.3-t2v

minimax/hailuo-2.3-t2v

text-to-video

[Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.

Text to Speech

minimax/speech-2.8-turbo

minimax/speech-2.8-turbo

text-to-speech

[Core Function] MiniMax Speech 2.8 Turbo is a lower-latency text-to-speech model with the same control surface as Speech 2.8 HD, including paralinguistic tags such as (laughs). [Strengths] Faster and more cost-efficient synthesis while retaining prosody, timbre mix, pronunciation, subtitle, and audio-format controls. [Best For] Highly recommended for: interactive assistants, high-volume TTS batches, cost-sensitive voiceovers, quick narration drafts, and latency-sensitive product prompts. [Limitations] Do NOT use this for streaming or real-time partial audio. Do NOT request hex output or emotion whisper. Do NOT send both voice_id and timbre_weights. Prefer speech-2.8-hd if maximum audio quality is required. Audio URLs expire in about 24 hours. [Routing] Choose speech-2.8-turbo when the user emphasizes speed, cost, or throughput. Choose speech-2.8-hd for premium quality or highly expressive delivery.

minimax/speech-2.8-hd

minimax/speech-2.8-hd

text-to-speech

[Core Function] MiniMax Speech 2.8 HD is a high-quality text-to-speech model that converts text into natural spoken audio, including expressive paralinguistic cues such as (laughs) and (sighs). [Strengths] Strong narration quality, stable prosody controls (speed, volume, pitch, emotion), optional timbre mixing, pronunciation overrides, subtitles, and flexible audio formats. [Best For] Highly recommended for: brand voiceovers, audiobook or long-form narration, marketing clips, multilingual delivery with language_boost, and expressive character speech. [Limitations] Do NOT use this for streaming or real-time partial audio. Do NOT request hex output or emotion whisper. Do NOT send both voice_id and timbre_weights. Audio URLs expire in about 24 hours. [Routing] Choose speech-2.8-hd when quality or expressive delivery matters most. Choose speech-2.8-turbo when the user prioritizes lower latency or cost.

Explore More

Image to Video
Category56 models
Text to Video
Category25 models
Video to Video
Category27 models
Wan
Wan
Series11 models
Gemini Omni
Gemini Omni
Series4 models
AI Animation Generator
AI Animation Generator
Collection16 models