
MiniMax H3モデルについて
MiniMax H3 is an open-weight omni-modal video model that reads text, images, video, and audio as one context and outputs up to 15 seconds of 2K video with native stereo audio.
すべてのMiniMax H3モデル

minimax/minimax-h3-max-v2v
[Core Function] MiniMax H3 Max V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 5-15 second duration, 480P or 768P output, optional aspect ratio defaulting to 16:9, and optional prompt_expansion_mode (disabled, balanced, or quality). [Best For] Highly recommended for: motion remix, style transfer from video clips, keeping subject motion while changing the scene description, and multi-clip reference guidance. [Limitations] Do NOT use this model for text-only generation, first or last frame transitions, or if you need 2K or a 4 second clip. prompt, duration, resolution, and reference_videos are required. Do NOT send first_frame_image or last_frame_image on this endpoint. [Routing] Choose MiniMax H3 Max V2V when the user provides reference videos. For reference stills only, use MiniMax H3 Max I2V. For start or end frames, use MiniMax H3 Max FL2V. For text only, use MiniMax H3 Max T2V. For 2K or 4 second clips, use MiniMax H3 V2V.

minimax/minimax-h3-max-i2v
[Core Function] MiniMax H3 Max I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 5-15 second duration, 480P or 768P output, optional aspect ratio defaulting to 16:9, and optional prompt_expansion_mode (disabled, balanced, or quality). [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions, when the primary input is a reference video, or if you need 2K or a 4 second clip. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 Max I2V when the user provides reference stills. For start or end frames, use MiniMax H3 Max FL2V. For reference videos, use MiniMax H3 Max V2V. For text only, use MiniMax H3 Max T2V. For 2K or 4 second clips, use MiniMax H3 I2V.

minimax/minimax-h3-max-fl2v
[Core Function] MiniMax H3 Max FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 5-15 second duration and 480P or 768P output. Output framing follows the input frame imagery. Optional prompt_expansion_mode is disabled, balanced (default), or quality. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix, or if you need 2K or a 4 second clip. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send ratio, reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 Max FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 Max I2V. For reference videos, use MiniMax H3 Max V2V. For text only, use MiniMax H3 Max T2V. For 2K or 4 second clips, use MiniMax H3 FL2V.

minimax/minimax-h3-max-t2v
[Core Function] MiniMax H3 Max T2V creates video from a text prompt only. [Strengths] It supports 5-15 second clips, 480P or 768P output, concrete aspect ratios from cinematic ultrawide to vertical, and optional prompt_expansion_mode (disabled, balanced, or quality). [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos, or if you need 2K or a 4 second clip. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 Max T2V for text-only MiniMax H3 Max video. If the user provides a start or end frame, use MiniMax H3 Max FL2V. If they provide reference images, use MiniMax H3 Max I2V. If they provide reference videos, use MiniMax H3 Max V2V. For 2K or 4 second clips, use MiniMax H3 T2V.

minimax/minimax-h3-v2v
[Core Function] MiniMax H3 V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: motion remix, style transfer from video clips, keeping subject motion while changing the scene description, and multi-clip reference guidance. [Limitations] Do NOT use this model for text-only generation or first or last frame transitions. prompt, duration, resolution, and reference_videos are required. Do NOT send first_frame_image or last_frame_image on this endpoint. [Routing] Choose MiniMax H3 V2V when the user provides reference videos. For reference stills only, use MiniMax H3 I2V. For start or end frames, use MiniMax H3 FL2V. For text only, use MiniMax H3 T2V.

minimax/minimax-h3-i2v
[Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions or when the primary input is a reference video. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 I2V when the user provides reference stills. For start or end frames, use MiniMax H3 FL2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

minimax/minimax-h3-fl2v
[Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 I2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

minimax/minimax-h3-t2v
[Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.


