minimax/minimax-h3-max-i2v

minimax-h3-max-i2v
Docs
Schema

[Core Function] MiniMax H3 Max I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 5-15 second duration, 480P or 768P output, optional aspect ratio defaulting to 16:9, and optional prompt_expansion_mode (disabled, balanced, or quality). [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions, when the primary input is a reference video, or if you need 2K or a 4 second clip. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 Max I2V when the user provides reference stills. For start or end frames, use MiniMax H3 Max FL2V. For reference videos, use MiniMax H3 Max V2V. For text only, use MiniMax H3 Max T2V. For 2K or 4 second clips, use MiniMax H3 I2V.

$0.0500~$0.0800/sec
image-to-video

Input

Video content description.
Required video duration in seconds. 4 seconds is not supported.
Reference image URLs used to guide subject or style (1-9).
Hint: Drag and drop files, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Required output resolution. 2K is not supported.
Optional reference audio URLs (up to 3).
Optional prompt expansion. disabled turns expansion off; balanced is the default; quality prioritizes expansion quality.
Optional video aspect ratio. One of the listed values. Default: 16:9.

Result

No results yet

Run the model to preview the output here.

Next:

README

Attributes

Series

Parameters

Name Description Type Required Enums
prompt Video content description. string Yes -
duration Required video duration in seconds. 4 seconds is not supported. integer Yes 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
reference_images Reference image URLs used to guide subject or style (1-9). string[] Yes -
resolution Required output resolution. 2K is not supported. string Yes 480P, 768P
reference_audios Optional reference audio URLs (up to 3). string[] No -
prompt_expansion_mode Optional prompt expansion. disabled turns expansion off; balanced is the default; quality prioritizes expansion quality. string No disabled, balanced, quality
ratio Optional video aspect ratio. One of the listed values. Default: 16:9. string No 21:9, 16:9, 4:3, 1:1, 3:4, 9:16

Pricing

Unit: $/sec

Dimension Pricing
resolution: 480P 0.0500
resolution: 768P 0.0800
  • minimax/minimax-h3-max-t2v: [Core Function] MiniMax H3 Max T2V creates video from a text prompt only. [Strengths] It supports 5-15 second clips, 480P or 768P output, concrete aspect ratios from cinematic ultrawide to vertical, and optional prompt_expansion_mode (disabled, balanced, or quality). [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos, or if you need 2K or a 4 second clip. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 Max T2V for text-only MiniMax H3 Max video. If the user provides a start or end frame, use MiniMax H3 Max FL2V. If they provide reference images, use MiniMax H3 Max I2V. If they provide reference videos, use MiniMax H3 Max V2V. For 2K or 4 second clips, use MiniMax H3 T2V.
  • minimax/minimax-h3-t2v: [Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.
  • minimax/minimax-h3-fl2v: [Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 I2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.
  • minimax/minimax-h3-i2v: [Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions or when the primary input is a reference video. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 I2V when the user provides reference stills. For start or end frames, use MiniMax H3 FL2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.
  • minimax/minimax-h3-max-fl2v: [Core Function] MiniMax H3 Max FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 5-15 second duration and 480P or 768P output. Output framing follows the input frame imagery. Optional prompt_expansion_mode is disabled, balanced (default), or quality. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix, or if you need 2K or a 4 second clip. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send ratio, reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 Max FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 Max I2V. For reference videos, use MiniMax H3 Max V2V. For text only, use MiniMax H3 Max T2V. For 2K or 4 second clips, use MiniMax H3 FL2V.
  • minimax/minimax-h3-v2v: [Core Function] MiniMax H3 V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: motion remix, style transfer from video clips, keeping subject motion while changing the scene description, and multi-clip reference guidance. [Limitations] Do NOT use this model for text-only generation or first or last frame transitions. prompt, duration, resolution, and reference_videos are required. Do NOT send first_frame_image or last_frame_image on this endpoint. [Routing] Choose MiniMax H3 V2V when the user provides reference videos. For reference stills only, use MiniMax H3 I2V. For start or end frames, use MiniMax H3 FL2V. For text only, use MiniMax H3 T2V.
  • minimax/minimax-h3-max-v2v: [Core Function] MiniMax H3 Max V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 5-15 second duration, 480P or 768P output, optional aspect ratio defaulting to 16:9, and optional prompt_expansion_mode (disabled, balanced, or quality). [Best For] Highly recommended for: motion remix, style transfer from video clips, keeping subject motion while changing the scene description, and multi-clip reference guidance. [Limitations] Do NOT use this model for text-only generation, first or last frame transitions, or if you need 2K or a 4 second clip. prompt, duration, resolution, and reference_videos are required. Do NOT send first_frame_image or last_frame_image on this endpoint. [Routing] Choose MiniMax H3 Max V2V when the user provides reference videos. For reference stills only, use MiniMax H3 Max I2V. For start or end frames, use MiniMax H3 Max FL2V. For text only, use MiniMax H3 Max T2V. For 2K or 4 second clips, use MiniMax H3 V2V.

Related Resources