minimax/hailuo-2.3-i2v

hailuo-2.3-i2v
文档
Schema

[Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.

$0.0540~$0.1200/sec
image-to-video

输入

Video content description, supports Chinese and English. Supports 15 camera control instructions: [Truck left/right], [Pan left/right], [Push in/Pull out], [Pedestal up/down], [Tilt up/down], [Zoom in/out], [Shake], [Tracking shot], [Static shot]. Use combined commands like [Pan left,Pedestal up] or sequential commands
First frame image URL or Base64 Data URL. Formats: JPG/JPEG/PNG/WebP. Size: <20MB. Dimensions: short side >300px, aspect ratio between 2:5 and 5:2
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Video duration in seconds. Note: 10 seconds only supports 768P resolution
6
Fast pretreatment mode (only effective when prompt_optimizer=true). Speeds up processing with slight quality trade-off
Enable automatic prompt optimization to improve video quality
Video resolution. Note: 10 seconds duration only supports 768P, 1080P only supports 6 seconds
768P

结果

暂无结果

运行模型后,结果将在这里显示。

Next:

README

参数

参数名 描述 类型 必填 枚举值
prompt Video content description, supports Chinese and English. Supports 15 camera control instructions: [Truck left/right], [Pan left/right], [Push in/Pull out], [Pedestal up/down], [Tilt up/down], [Zoom in/out], [Shake], [Tracking shot], [Static shot]. Use combined commands like [Pan left,Pedestal up] or sequential commands string -
first_frame_image First frame image URL or Base64 Data URL. Formats: JPG/JPEG/PNG/WebP. Size: <20MB. Dimensions: short side >300px, aspect ratio between 2:5 and 5:2 string -
duration Video duration in seconds. Note: 10 seconds only supports 768P resolution integer 6, 10
fast_pretreatment Fast pretreatment mode (only effective when prompt_optimizer=true). Speeds up processing with slight quality trade-off boolean true, false
prompt_optimizer Enable automatic prompt optimization to improve video quality boolean true, false
resolution Video resolution. Note: 10 seconds duration only supports 768P, 1080P only supports 6 seconds string 768P, 1080P

价格

单位: $/sec

维度 价格
duration: 10 / resolution: 1080P 0.1200
duration: 10 / resolution: 768P 0.0672
duration: 6 / resolution: 1080P 0.1060
duration: 6 / resolution: 768P 0.0540

相关模型

  • minimax/hailuo-2.3-fast-i2v: [Core Function] Hailuo 2.3 Fast I2V is a high-speed, cost-effective image-to-video generation model. [Strengths] It excels at generating videos from images much faster and at roughly 50% lower cost than the standard 2.3 model, while still maintaining the 2.3 architecture’s strength in human motion. [Best For] Highly recommended for: rapid prototyping, batch social media creation, and cost-sensitive video generation pipelines. [Limitations] Do NOT use this model for text-to-video (it only accepts image inputs). Do NOT use when absolute maximum visual fidelity is the primary requirement. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or ‘cost-effective’ image-to-video generation. For maximum quality, use the standard Hailuo 2.3 I2V.
  • minimax/hailuo-2.3-t2v: [Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.
  • google/nano-banana-2: [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism (use the Imagen 4 series instead). [Routing] Route to this model for ‘fast’, ‘creative’, or ‘stylized’ high-volume requests.
  • google/nano-banana-pro: [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [Limitations] Do NOT use for absolute photorealism (use Imagen 4). [Routing] Use this model for high-quality, creative, non-photorealistic requests.
  • kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image generation model. [Strengths] It delivers state-of-the-art aesthetic quality, high-resolution outputs (up to 2K), and exceptional prompt adherence. [Best For] Highly recommended for: professional concept art, photorealistic portraits, and high-fidelity image generation. [Limitations] Do NOT use this model if you need to strictly reference or edit an existing image. [Routing] Use this as the default model for all text-to-image requests on the Kling platform.
  • minimax/minimax-image-01-t2i: [Core Function] MiniMax Image-01 T2I is a multimodal text-to-image generation model. [Strengths] It excels at blending high-quality image generation with visual reasoning, allowing for strong prompt adherence and structural understanding. [Best For] Highly recommended for: general image generation, conceptual illustrations, and generating multiple images in a single batch. [Limitations] Do NOT use this model if the user specifically requests video generation. [Routing] Use this as the default text-to-image model for MiniMax API integrations.
  • vidu/viduq3-pro-t2v: [Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot ‘smart cuts’, complex physics, and simultaneous audio-visual generation. [Best For] Highly recommended for: cinematic storytelling, professional advertising, short films, and high-fidelity concept visualizations. [Limitations] Do NOT use this model if the user is looking for an instant, low-latency preview, as generation takes longer. It does not support automatic BGM addition. [Routing] Use this model by default for high-quality text-to-video requests. If the user specifically asks for ‘fast’ or ‘quick’ generation, switch to the Q3 Turbo T2V model.
  • vidu/viduq3-turbo-t2v: [Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, generating background b-roll, and scenarios where generation speed is prioritized. [Limitations] Do NOT use this model if you need ultimate cinematic quality, complex audio-visual synchronization, or multi-shot ‘smart cuts’. It does not support automatic BGM addition. [Routing] Choose this ‘Turbo’ model when the user emphasizes ‘quick’, ‘fast’, or needs immediate results. If the user demands the highest cinematic quality or advanced audio-visual features, choose the Q3 Pro T2V model instead.
  • vidu/viduq3-r2v: [Core Function] Vidu Q3 R2V is a high-quality reference-to-video generation model. [Strengths] It excels at generating detailed, cinematic videos that precisely follow a text prompt while highly preserving the character identity from provided reference images. [Best For] Highly recommended for: professional character-driven storytelling, high-fidelity avatar generation in new scenes, and cinematic films requiring consistent actors. [Limitations] Do NOT use this model if you just want to add motion to an existing image (use I2V). This model creates new scenes based on the prompt while keeping the character. [Routing] Use this by default when the user wants to generate a video of a specific character (provided via image) doing something new (provided via text prompt).
  • pixverse/video-restyle: [Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it only re-renders the visual style of an existing video. It requires an input video, and you must provide EITHER restyle_id (a preset style code from the PixVerse restyle list) OR restyle_prompt (free-text style, max 2048 chars), not both. [Routing] Use when the user wants to change the look of an existing video. Use restyle_id for an official preset, restyle_prompt for a custom style.
  • kling/kling-v3-i2v: [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.
  • bytedance/seedance-2.0-i2v: [Core Function] Seedance 2.0 I2V is ByteDance’s flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.
  • google/veo-3.1-i2v: [Core Function] Veo 3.1 I2V is Google’s cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animating concept art, creating cinematic transitions between images, and high-end video production. [Limitations] When using first+last frame or reference-only modes, the duration parameter must strictly be 8. Negative prompts are not supported in reference-only mode. [Routing] Use this model by default for high-quality image-to-video tasks or when multiple reference images are provided.
  • vidu/viduq3-pro-i2v: [Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended for: bringing concept art to life, professional film production, high-end commercial showcases, and creating immersive environments from still images. [Limitations] Do NOT use this model if you need instant/real-time generation, as rendering takes longer. It does not support 4K resolution. [Routing] Use this model by default for high-quality image-to-video requests. If the user requires faster generation, route to Q3 Pro Fast or Q3 Turbo.
  • vidu/viduq3-pro-fl2v: [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.
  • skywork/skyreels-i2v: [Core Function] SkyReels Image-to-Video animates one or more keyframe images into a video guided by a text prompt. [Strengths] Supports a start frame, an end frame, and tagged mid-frames for keyframe control; optional audio, 480p/720p/1080p output, and fast/std modes. [Best For] Bringing a photo to life, first-last-frame transitions, and keyframe-driven storyboards. [Limitations] Do NOT use this for pure text-to-video (use skyreels-t2v) or for editing an existing video (use the Omni / video-to-video models). It requires at least one of first_frame_image, end_frame_image, or mid_frame_images; output is capped at 1080p and 15s, and fast mode supports only sound=false. [Routing] Provide first_frame_image to animate from a start image, add end_frame_image for a transition, or supply mid_frame_images (each tag must appear in the prompt as @tag) for keyframe guidance.
  • pixverse/motion-control: [Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limitations] Do NOT use this if you only have a video and no subject image (use Restyle or Extend instead), or if you need 1080p output (only 360p/540p/720p are supported). It requires BOTH a subject image (with a clear person or animal) AND a reference video (with a person as the primary focus). [Routing] Use when the user has one subject image and one motion reference video and wants the subject to mimic that motion.
  • vidu/motion-sync: [Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target’s identity. [Best For] Highly recommended for: creating dance videos with custom characters, transferring complex choreography, and replicating specific physical actions onto avatars. [Limitations] Do NOT use this model if you want to change what a character is saying (use Lip Sync). It requires both a reference video for motion and a target image for appearance. [Routing] Use this model specifically when the user wants to copy the body movements or actions from one video onto a different character.

同系列模型