Explore 77 production-ready Image to Video AI models on Modellix. Compare capabilities, try models in the playground, and integrate them through one unified API.

[Core Function] Gemini Omni Flash R2V (Reference-to-Video) generates a short 720p video guided by up to three reference images via the Interactions API. [Strengths] It fuses the styles, subjects, or elements from multiple reference images (referred to in the text prompt) into a single coherent animated clip with synchronized audio. [Best For] Highly recommended for: blending characters or visual styles from several images, reference-guided creative shots, and multi-subject compositions where the prompt directs how the references combine. [Limitations] Do NOT use this model if you only have a single starting frame (use I2V instead), or if you need 1080p or 4K or clips longer than 10 seconds; it accepts 1 to 3 reference images and outputs 720p up to 10 seconds (16:9 or 9:16). [Routing] Choose this when the user supplies multiple reference images to combine into one video. For single first-frame animation use Gemini Omni Flash I2V; to modify an existing video use Gemini Omni Flash Video Edit.

[Core Function] Gemini Omni Flash I2V is a fast Image-to-Video model that animates a single input image into a short 720p video via the Interactions API. [Strengths] It uses the provided image as the opening frame and generates smooth motion with natively synchronized audio at low latency. [Best For] Highly recommended for: bringing a still photo to life, quick product or portrait animation, and short social clips derived from a single image. [Limitations] Do NOT use this model if you need 1080p or 4K output, clips longer than 10 seconds, or the fusion of multiple reference images; it takes exactly one image and outputs 720p up to 10 seconds (16:9 or 9:16). Do NOT use it to edit an existing video. [Routing] Choose this when the user provides one image to animate. To fuse multiple reference images use Gemini Omni Flash R2V; to edit an existing video use Gemini Omni Flash Video Edit; for 4K cinematic results use Veo 3.1 I2V.

**[Core Function]** SkyReels Segmented Camera Motion (audio-to-video) generates a talking-avatar video with directed camera movement across time segments. **[Strengths]** Combines an audio-driven avatar with per-segment camera trajectories such as push, pan, crane and rotation. **[Best For]** Dynamic presenter clips and cinematic avatar shots with controlled camera motion. **[Limitations]** Do NOT use this when you need a completely static camera (use single-actor-avatar). It requires first_frame_image and one audio segment; use camera_control_pro for multi-segment or compound moves. mode=std outputs 720p, mode=pro outputs 1080p. **[Routing]** Set a single traj_type plus camera_control_strength for a simple move, or supply camera_control_pro (a list of per-segment {start_time, end_time, traj_type, ...}) for compound motion; choose mode=pro for 1080p.

**[Core Function]** SkyReels Single-Actor Avatar (audio-to-video) drives a talking-avatar video from a single portrait image and one audio track. **[Strengths]** Lip-synced single-speaker talking-head video generated from an image plus audio. **[Best For]** Virtual presenters, single-speaker dubbing, and talking avatars. **[Limitations]** Do NOT use this for multi-speaker scenes (use the multi-actor flow) or when you only have text. It requires first_frame_image and exactly one audio segment (<=200s). mode=std outputs 720p, mode=pro outputs 1080p. **[Routing]** Provide a portrait first_frame_image and one audio URL in audios; choose mode=pro for 1080p output.

**[Core Function]** SkyReels Reference-to-Video (multiobject) generates a video from a prompt while preserving the subjects from 1-4 reference images. **[Strengths]** Multi-subject identity preservation, placing specific characters or objects into a newly generated scene. **[Best For]** Putting given characters/products into a new scene, multi-subject composition from reference photos. **[Limitations]** Do NOT use this to animate a single fixed frame (use skyreels-i2v) or for text-only generation (use skyreels-t2v). It requires 1-4 reference images and produces clips up to 5s. **[Routing]** Provide 1-4 subject reference images in ref_images; set aspect_ratio and duration (1-5s) as needed.

**[Core Function]** SkyReels Image-to-Video animates one or more keyframe images into a video guided by a text prompt. **[Strengths]** Supports a start frame, an end frame, and tagged mid-frames for keyframe control; optional audio, 480p/720p/1080p output, and fast/std modes. **[Best For]** Bringing a photo to life, first-last-frame transitions, and keyframe-driven storyboards. **[Limitations]** Do NOT use this for pure text-to-video (use skyreels-t2v) or for editing an existing video (use the Omni / video-to-video models). It requires at least one of first_frame_image, end_frame_image, or mid_frame_images; output is capped at 1080p and 15s, and fast mode supports only sound=false. **[Routing]** Provide first_frame_image to animate from a start image, add end_frame_image for a transition, or supply mid_frame_images (each tag must appear in the prompt as @tag) for keyframe guidance.

[Core Function] PixVerse c1 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composition and identity preservation. [Best For] Putting specific characters/objects into a new scene, multi-character interactions. [Limitations] Do NOT use this for single-image animation (use Image-to-Video), two-frame transitions (use First-Last-Frame), or text-only generation (use Text-to-Video). It requires 1-7 reference images and a required aspect_ratio, and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one or more reference images to be fused into the video, especially when distinguishing subject vs background or naming references for the prompt. The v6 and c1 reference-to-video variants accept identical parameters; choose the version the user requests.

[Core Function] PixVerse v6 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composition and identity preservation. [Best For] Putting specific characters/objects into a new scene, multi-character interactions. [Limitations] Do NOT use this for single-image animation (use Image-to-Video), two-frame transitions (use First-Last-Frame), or text-only generation (use Text-to-Video). It requires 1-7 reference images and a required aspect_ratio, and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one or more reference images to be fused into the video, especially when distinguishing subject vs background or naming references for the prompt. The v6 and c1 reference-to-video variants accept identical parameters; choose the version the user requests.

[Core Function] PixVerse c1 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitations] Do NOT use this if you only have a single image (use Image-to-Video) or lack an end frame; it requires exactly two images (a first and a last frame) and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides a start image AND an end image. For one image use Image-to-Video. The v6 and c1 first-last-frame variants accept identical parameters; choose the version the user requests.

[Core Function] PixVerse v6 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitations] Do NOT use this if you only have a single image (use Image-to-Video) or lack an end frame; it requires exactly two images (a first and a last frame) and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides a start image AND an end image. For one image use Image-to-Video. The v6 and c1 first-last-frame variants accept identical parameters; choose the version the user requests.

[Core Function] PixVerse c1 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic motion from a still. [Limitations] Do NOT use this for two-frame transitions (use First-Last-Frame), multi-subject fusion (use Reference-to-Video), or text-only generation (use Text-to-Video). It requires exactly one starting image, does not support multi-clip, and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one starting image. The c1 variant matches v6 inputs but does not support multi-clip generation; choose c1-i2v when the user requests the c1 model and multi-clip is not needed.

[Core Function] PixVerse v6 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio and multi-clip. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic motion from a still. [Limitations] Do NOT use this for two-frame transitions (use First-Last-Frame), multi-subject fusion (use Reference-to-Video), or text-only generation (use Text-to-Video). It requires exactly one starting image and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one starting image. The v6 and c1 variants take the same inputs except v6 also supports multi-clip generation (generate_multi_clip_switch); choose v6-i2v when the user wants multi-clip output or requests the v6 model.

[Core Function] Vidu Q3 AD is a keyframe-driven short-play (short drama) Image-to-Video model that turns a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass. [Strengths] It excels at multi-shot narrative coherence and keeping referenced characters, scenes, and props consistent across shots, driven directly from keyframe reference images without manual shot-by-shot prompting. [Best For] Highly recommended for: short ad films and brand stories, scripted short dramas, multi-scene narrative clips generated from a script, and character-driven reels built from a small cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image (use a standard image-to-video model such as Vidu Q3 Pro/Turbo I2V instead); it requires script_content in traditional screenplay format (scene + characters + dialogue) plus 1-14 reference assets each with an image_uri, outputs 1080p in 16:9 or 9:16, and does not expose per-shot duration or style controls (duration is auto-planned). Asset type must be one of character/scene/tool. [Routing] Choose Vidu Q3 AD when the user provides a script and reference images and wants an automatically directed multi-shot short drama from keyframes. For animating a single image into one continuous clip, route to viduq3-pro-i2v / viduq3-turbo-i2v; for the standard short-play flow with placement/quality/duration/style controls, route to Vidu Q3 Drama.

[Core Function] Seedance 2.0 Mini I2V is the lightweight multimodal image-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with optional first/last frame, reference images, and audio references. [Routing] Use for cost-efficient image-to-video.

[Core Function] HappyHorse 1.1 R2V is Alibaba's latest reference-image-to-video model. [Strengths] It uses 1-9 reference images to preserve subject or character appearance while generating new video actions, supports 720P/1080P output, 3-15 second duration, and expanded aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: character-consistent storytelling, reference-based product shots, and multi-image subject composition. [Limitations] Do NOT use if the user simply wants to animate a single image exactly as provided; use HappyHorse I2V instead. [Routing] Use when the user provides one or more reference images and asks for a newly generated HappyHorse video.

[Core Function] HappyHorse 1.1 I2V is Alibaba's latest streamlined first-frame image-to-video model. [Strengths] It turns a single image into high-quality 720P/1080P video with native audio support and 3-15 second duration; output aspect ratio follows the first frame image. [Best For] Highly recommended for: rapid image animation, product motion previews, and simple character or scene animation. [Limitations] It does not accept an explicit ratio parameter; use T2V or R2V when you need a fixed generated aspect ratio. [Routing] Prefer this model when the user provides one image and requests HappyHorse image animation.

[Core Function] Grok Imagine Video 1.5 I2V animates a single starting image into a video using the Grok Imagine 1.5 generation backbone. [Strengths] It excels at producing motion from one starting frame with the improved 1.5 model. [Best For] Highly recommended for: animating a photo or illustration when the 1.5 generation backbone is preferred. [Limitations] Do NOT use this model for text-only generation, for resolutions above 1080p, or for clips longer than 15 seconds; it requires a starting image and is limited to 480p/720p/1080p and 15s. [Routing] Use this model when the user provides one starting image and prefers the 1.5 backbone.

[Core Function] Grok Imagine Video R2V generates a video from a text prompt while preserving the subjects shown in up to 7 reference images. [Strengths] It excels at keeping character/subject identity consistent across a newly generated scene driven by the prompt. [Best For] Highly recommended for: character-driven video, placing a specific subject into a new scene, and blending features from multiple references. [Limitations] Do NOT use this model to simply animate a single image as-is (use Image-to-Video), or for clips longer than 10 seconds or resolutions above 720p; duration is capped at 10s and resolution at 480p/720p. [Routing] Use this model when the user provides reference images of a subject and wants a new action/scene described by a prompt. To simply animate a single image as-is, use Image-to-Video.

[Core Function] Grok Imagine Video I2V animates a single starting image into a video. [Strengths] It excels at producing smooth motion from one starting frame, guided by a text prompt for the desired movement. [Best For] Highly recommended for: bringing a photo or illustration to life, dynamic product showcases, and quick cinematic motion from a still. [Limitations] Do NOT use this model for text-only generation, for resolutions above 720p, or for clips longer than 15 seconds; it requires a starting image and is limited to 480p/720p and 15s. [Routing] Use this model when the user provides exactly one starting image. For reference-driven character video, use Reference-to-Video; for text-only video, use Text-to-Video.

[Core Function] Vidu Q3 Drama (Short Play) is a script-to-video model that turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass. [Strengths] It excels at multi-shot narrative coherence, automatic storyboarding and cinematography, and keeping the identity of referenced characters, scenes, and props consistent across every shot. [Best For] Highly recommended for: scripted short dramas and web-series episodes, narrative short-form ads and brand stories, rapid storyboard and pre-visualization, and character-driven reels built from a cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image as-is (use a standard image-to-video model such as Vidu Q3 Pro instead); it requires a script and 1-14 reference assets, supports only 8-12 second clips at 1080p in 16:9 or 9:16, and is not intended for pixel-perfect single-product shots or complex multi-object physics. [Routing] Choose Vidu Q3 Drama when the user provides a script or narrative beats plus character/scene/prop references and wants an automatically directed multi-shot short play. If the user only wants to animate a single image or needs one continuous clip without scripted scene changes, route to a standard image-to-video model such as Vidu Q3 Pro.

[Core Function] Vidu Q3 Turbo R2V is a fast reference-to-video generation model. [Strengths] It excels at quickly generating dynamic videos based on a text prompt while preserving the identity of the subjects from provided reference images. [Best For] Highly recommended for: rapid character animation, quick conceptual mockups with specific subjects, and fast social media content featuring consistent characters. [Limitations] Do NOT use this model if you need the absolute highest cinematic quality or if you just want to animate an image directly without a text prompt. [Routing] Choose this model for fast generation of a character performing actions based on a text prompt. For better quality, use Q3 R2V or Q3 Mix R2V.

[Core Function] Vidu Q3 R2V is a high-quality reference-to-video generation model. [Strengths] It excels at generating detailed, cinematic videos that precisely follow a text prompt while highly preserving the character identity from provided reference images. [Best For] Highly recommended for: professional character-driven storytelling, high-fidelity avatar generation in new scenes, and cinematic films requiring consistent actors. [Limitations] Do NOT use this model if you just want to add motion to an existing image (use I2V). This model creates new scenes based on the prompt while keeping the character. [Routing] Use this by default when the user wants to generate a video of a specific character (provided via image) doing something new (provided via text prompt).

[Core Function] Vidu Q3 Mix R2V is a mixed-style reference-to-video generation model. [Strengths] It excels at generating highly consistent character videos by synthesizing and blending features from multiple reference images (up to 7) based on a text prompt. [Best For] Highly recommended for: maintaining strict character consistency across different styles, generating videos of a specific subject in entirely new environments, and blending concepts from multiple reference images. [Limitations] Do NOT use this model if you just want to animate a single image as is (use standard I2V). This model focuses on extracting character/style features and generating new content. [Routing] Use this when the user provides reference images of a character/subject and wants a video of them doing a specific new action from a text prompt, prioritizing mixed-style consistency.

[Core Function] Vidu One-Click AD-Film is an automated marketing video generation model. [Strengths] It excels at transforming 1 to 7 product or scene images into a polished, commercial-style advertisement video (10-60s) automatically. [Best For] Highly recommended for: e-commerce product showcases, social media ads, promotional reels, and quick marketing campaigns. [Limitations] Do NOT use this model for narrative storytelling or cinematic films; it is optimized specifically for commercial pacing and product emphasis. [Routing] Use this specifically when the user wants to generate an 'ad', 'commercial', or 'promotional video' from product photos.

[Core Function] Vidu One-Click General Film is an automated cinematic film generation model. [Strengths] It excels at automatically stringing together 1 to 7 user-provided images into a cohesive, cinematic film (up to 180s) with appropriate transitions and pacing. [Best For] Highly recommended for: instant music videos, cinematic montages, automated travel vlogs, and turning photo albums into compelling short films. [Limitations] Do NOT use this model if the user needs precise, frame-by-frame control over camera movements or specific character actions in each shot. [Routing] Use this when the user wants an automated 'done-for-you' long video from a batch of images without manually prompting every single shot.

[Core Function] Vidu Q2 Pro Multi-Frame is a sequence-based animation model. [Strengths] It excels at creating continuous, high-quality animation by interpolating through a provided sequence of keyframes (up to 9 images). [Best For] Highly recommended for: complex motion control, precise character posing sequences, and animating detailed storyboards where intermediate states must be strictly followed. [Limitations] Do NOT use this model if you only have one or two images (use standard I2V or FL2V instead). It requires a start image and a list of key images. [Routing] Use this exclusively when the user provides a sequence of 3 to 9 specific frames and wants them animated in order.

Vidu Q2 Turbo multi-frame animation model. Animates through a sequence of up to 9 frames (1 start + up to 8 key images). Both `start_image` and `key_images` are required. Supports `resolution`.

[Core Function] Vidu Q2 Pro Digital Human is a premium portrait animation model. [Strengths] It excels at generating highly realistic, expressive digital humans from a single portrait image, featuring precise lip-sync to audio and natural facial micro-expressions. [Best For] Highly recommended for: professional virtual spokespersons, high-end educational videos, news anchoring, and realistic character animation. [Limitations] Do NOT use this model for complex full-body physical interactions or videos longer than 10 seconds. [Routing] Use this by default for 'talking head' or 'digital human' requests prioritizing realism over speed.

[Core Function] Vidu Q2 Turbo Digital Human is a fast portrait animation model. [Strengths] It excels at quickly animating a static portrait image into a speaking or moving digital human, syncing lip movements to provided audio with low latency. [Best For] Highly recommended for: rapid generation of talking head videos, quick virtual presenters, and responsive interactive avatars. [Limitations] Do NOT use this model for complex full-body motion, multi-character interactions, or videos longer than 10 seconds. [Routing] Use this model when the user wants to make a portrait 'talk' quickly. For higher realism and better quality, route to Q2 Pro Digital Human.

[Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.

[Core Function] Vidu Q3 Turbo FL2V is a fast First-Last frame transition video model. [Strengths] It excels at rapidly generating a smooth video transition bridging a specific starting image and an ending image. [Best For] Highly recommended for: quick visual morphs, before-and-after transitions, time-lapse simulations, and rapid storyboard filling. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. Do NOT use if you need the highest possible cinematic detail. [Routing] Use this when the user provides exactly two images and wants a fast transition between them. For higher quality transitions, use Q3 Pro FL2V.

[Core Function] Vidu Q3 Turbo I2V is a fast Image-to-Video generation model. [Strengths] It excels at quickly animating a starting frame into a video sequence with strong motion dynamics and low latency. [Best For] Highly recommended for: rapid social media content creation, quick animatics, and fast visual iterations from reference images. [Limitations] Do NOT use this model if you require the absolute highest cinematic fidelity or complex audio-visual synchronization. [Routing] Choose this model for general quick image-to-video tasks. For the highest quality, route to Q3 Pro I2V.

[Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended for: bringing concept art to life, professional film production, high-end commercial showcases, and creating immersive environments from still images. [Limitations] Do NOT use this model if you need instant/real-time generation, as rendering takes longer. It does not support 4K resolution. [Routing] Use this model by default for high-quality image-to-video requests. If the user requires faster generation, route to Q3 Pro Fast or Q3 Turbo.

[Core Function] Vidu Q3 Pro Fast I2V is a high-speed Image-to-Video generation model. [Strengths] It excels at generating smooth, physically accurate continuous motion from a single starting frame with extremely low latency. [Best For] Highly recommended for: fast prototyping, short dynamic product showcases, quick cinematic transitions, and scenarios where generation speed is prioritized over maximum detail. [Limitations] Do NOT use this model if the user requires 4K resolution, complex multi-character interactions, or highly stylized 2D anime deformations. It only supports 720p/1080p resolutions. [Routing] Choose this 'Fast' model when the user emphasizes 'quick', 'fast', or needs immediate results. If the user demands ultimate cinematic quality, choose the standard Q3 Pro I2V model instead.

[Core Function] Wan 2.7 I2V is Alibaba's flagship multimodal image-to-video model. [Strengths] It supports multimodal input (text, image, audio, video) for first-frame, start-and-end-frame (FL2V), and video continuation tasks. [Best For] Highly recommended for: complex image animation, cinematic transitions, and video extension workflows. [Limitations] Do NOT use this model if you only need a quick, simple animation where HappyHorse might be faster. [Routing] Use this model by default for complex image-to-video or video continuation tasks.

[Core Function] Veo 3.1 Lite I2V is a balanced image-to-video model. [Strengths] It offers a middle ground between speed and quality for animating images. [Best For] Highly recommended for: general image animation and web-ready content. [Limitations] Do NOT use this model if you need 4K resolution. [Routing] Use this model for standard image animation requests.

[Core Function] HappyHorse 1.0 R2V is a reference-to-video model. [Strengths] It excels at maintaining character consistency using up to 9 reference images while generating new video actions based on a prompt. [Best For] Highly recommended for: character-consistent storytelling and generating multiple scenes with the same subject. [Limitations] Do NOT use if you simply want to animate a single image exactly as it is (use HappyHorse I2V instead). [Routing] Use this when the user provides reference images to dictate character/subject appearance in a newly generated action.

[Core Function] HappyHorse 1.0 I2V is a streamlined image-to-video model. [Strengths] It generates high-quality 720P/1080P video (3-15s) from an image efficiently, with native audio support. [Best For] Highly recommended for: rapid image animation and robust character motion. [Limitations] Does not support complex video continuation like Wan 2.7 I2V. [Routing] Use when the user requests 'HappyHorse' or a streamlined image animation.

[Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.

[Core Function] Seedance 2.0 Fast I2V is a high-speed multimodal video generation model. [Strengths] Fast generation with the multimodal and multi-shot capabilities of the Seedance 2.0 architecture. [Best For] Highly recommended for: rapid prototyping and quick multi-shot video creation. [Limitations] Do NOT use this model for the absolute highest cinematic fidelity (use the standard 2.0 model instead). [Routing] Choose this model when the user emphasizes 'fast' or 'quick' generation.

[Core Function] Seedance 2.0 I2V is ByteDance's flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.

[Core Function] MiniMax S2V-01 is a Subject-to-Video generation model. [Strengths] It excels at maintaining strict identity consistency of a specific subject (provided via reference images) while generating a video of that subject performing actions described in a text prompt. [Best For] Highly recommended for: character-consistent storytelling, generating multiple videos with the same actor, and brand mascot animation. [Limitations] Do NOT use this model if you just want to animate the provided image exactly as it is (use standard I2V). This model generates *new* scenes featuring the referenced subject. [Routing] Use this model when the user provides reference images of a character/subject and wants a video of them doing a specific new action based on a text prompt.

[Core Function] Hailuo 02 FL2V is a First-Last frame transition video model. [Strengths] It excels at generating a logical, physically accurate video transition that bridges a provided starting frame and an ending frame. [Best For] Highly recommended for: visual morphing, before-and-after transitions, and precise narrative storyboard completion. [Limitations] Do NOT use this model if you only have one image (use standard I2V instead). Note that Hailuo 2.3 does not support FL2V, so this is the primary transition model. [Routing] Use this model exclusively when the user provides BOTH a first frame and a last frame for a transition.

[Core Function] MiniMax I2V-01 is a legacy image-to-video model. [Strengths] Standard image-to-video generation maintained for backward compatibility. [Best For] Recommended only for: maintaining existing integrations. [Limitations] Do NOT use this model for new creations. It is superseded by the Hailuo series. [Routing] Only use this if the user explicitly requests the legacy 'I2V-01' model.

[Core Function] MiniMax I2V-01-Live is a legacy image-to-video model optimized for anime and stylized art. [Strengths] Originally designed for fast, stylized animation, particularly for 2D illustration and anime styles. [Best For] Recommended for: specific legacy integrations that relied on the 'Live' anime aesthetics. [Limitations] Do NOT use this model for general generation. The new Hailuo 2.3 model has fully absorbed and improved upon these stylization capabilities. [Routing] Only use this if the user explicitly requests the 'Live' model. For anime/stylized art, use Hailuo 2.3 I2V instead.

[Core Function] MiniMax I2V-01-Director is a legacy image-to-video model with explicit camera controls. [Strengths] It allows direct manipulation of virtual camera movements (pan, zoom, tilt) applied to a starting image. [Best For] Recommended for: legacy systems requiring explicit camera parameter inputs. [Limitations] Do NOT use this model for general video generation; it is an older architecture with lower physical realism than the Hailuo series. [Routing] Only use this if the user explicitly requests the 'Director' model or legacy 'I2V-01' generation. Otherwise, default to Hailuo 2.3 I2V.

[Core Function] Hailuo 02 I2V is an image-to-video model optimized for physical realism and sustained high resolution. [Strengths] It excels at animating broad scenes, maintaining complex physics, and supporting native 1080p generation for up to 10 seconds. [Best For] Highly recommended for: animating product photography, bringing landscape/nature photos to life, and generating physically accurate motion. [Limitations] Do NOT use this model for animating complex human facial micro-expressions or highly stylized anime art, where Hailuo 2.3 is superior. [Routing] Use this model when the user needs to animate a landscape/product, or explicitly requires 10 seconds of 1080p video. Otherwise, default to Hailuo 2.3 I2V.

[Core Function] Hailuo 2.3 Fast I2V is a high-speed, cost-effective image-to-video generation model. [Strengths] It excels at generating videos from images much faster and at roughly 50% lower cost than the standard 2.3 model, while still maintaining the 2.3 architecture's strength in human motion. [Best For] Highly recommended for: rapid prototyping, batch social media creation, and cost-sensitive video generation pipelines. [Limitations] Do NOT use this model for text-to-video (it only accepts image inputs). Do NOT use when absolute maximum visual fidelity is the primary requirement. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or 'cost-effective' image-to-video generation. For maximum quality, use the standard Hailuo 2.3 I2V.

[Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.

[Core Function] Veo 2 I2V is an older generation image-to-video model. [Strengths] Known for its distinctive cinematic style and fluid motion priors from the Veo 2 era. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Veo 3.1 I2V.

[Core Function] Veo 3 Fast I2V is an older generation image-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Veo 3.1 I2V.

[Core Function] Veo 3 I2V is an older generation image-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Veo 3.1 I2V.

[Core Function] Veo 3.1 Fast I2V is a high-speed image-to-video model. [Strengths] It quickly animates starting images at 1080p, optimized for low latency. [Best For] Highly recommended for: rapid prototyping and quick social media visual iterations. [Limitations] Do NOT use this model for the absolute highest visual fidelity or 4K output. [Routing] Choose this model when the user emphasizes 'fast' or 'quick' animation.

[Core Function] Veo 3.1 I2V is Google's cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animating concept art, creating cinematic transitions between images, and high-end video production. [Limitations] When using first+last frame or reference-only modes, the `duration` parameter must strictly be 8. Negative prompts are not supported in reference-only mode. [Routing] Use this model by default for high-quality image-to-video tasks or when multiple reference images are provided.

[Core Function] Kling Video Effects applies predefined visual effects to images. [Strengths] It automatically transforms 1 or 2 images into engaging short videos using viral/predefined effect templates. [Best For] Highly recommended for: social media trends and quick visual gags. [Limitations] Do NOT use this model for custom narrative generation or prompt-based control. [Routing] Use this when the user explicitly requests an 'effect' or 'trend' template applied to their photos.

[Core Function] Kling Avatar is a specialized portrait animation model. [Strengths] It precisely animates a portrait image to lip-sync with an audio file or TTS audio ID. [Best For] Highly recommended for: virtual presenters, talking head videos, and digital avatars. [Limitations] Do NOT use this model for full-body action or general image animation. [Routing] Use this model explicitly when the user wants to make a portrait 'speak' with provided audio.

[Core Function] Kling V1.6 MI2V is an older generation image-to-video model. [Strengths] Pioneered early realistic physics simulation in video generation. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 I2V.

[Core Function] Kling V2.6 I2V is a stable, classic image-to-video model. [Strengths] It provides excellent motion stability, motion brush controls, and 1080p generation with optional audio. [Best For] Highly recommended for: animating specific areas using motion brushes and stable 5s/10s generations. [Limitations] Do NOT use this model if you need 4K resolution or 15s durations. [Routing] Route to this model when users need Motion Brush capabilities or prefer the V2.6 aesthetic. Otherwise, use V3.

[Core Function] Kling V2.5 Turbo I2V is a fast, cost-effective image-to-video model. [Strengths] Fast generation speed and high value at 1080p. [Best For] Highly recommended for: rapid image animation and bulk processing. [Limitations] Do NOT use this model for the absolute highest fidelity or long 15s generations. [Routing] Choose this model when the user emphasizes speed or cost.

[Core Function] Kling V2.1 Master I2V is an older generation image-to-video model. [Strengths] Offered improved multi-subject tracking and better texture details over V1. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 I2V.

[Core Function] Kling V2.1 I2V is an older generation image-to-video model. [Strengths] Offered improved multi-subject tracking and better texture details over V1. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 I2V.

[Core Function] Kling V2 Master I2V is an older generation image-to-video model. [Strengths] Offered improved multi-subject tracking and better texture details over V1. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 I2V.

[Core Function] Kling V1.6 I2V is an older generation image-to-video model. [Strengths] Pioneered early realistic physics simulation in video generation. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 I2V.

[Core Function] Kling V1.5 I2V is an older generation image-to-video model. [Strengths] Pioneered early realistic physics simulation in video generation. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 I2V.

[Core Function] Kling V1 I2V is an older generation image-to-video model. [Strengths] Pioneered early realistic physics simulation in video generation. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 I2V.

[Core Function] Seedance 1.5 Pro I2V is a joint audio-video image-to-video model. [Strengths] It accurately follows complex instructions to animate a single image with synchronized audio. [Best For] Recommended for: standard image animation tasks where 2.0's multi-reference capabilities are not required. [Limitations] Lacks the robust multi-shot and multi-image reference features of 2.0. [Routing] Use this if the user specifically requests the 1.5 architecture, otherwise default to 2.0.

[Core Function] Seedance 1.0 Pro I2V is an older generation image-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Seedance 2.0 I2V.

[Core Function] Seedance 1.0 Pro Fast I2V is an older generation image-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Seedance 2.0 I2V.

[Core Function] Wanx 2.1 KF2V Plus is a legacy generation model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7.

[Core Function] Wan 2.2 KF2V Flash is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7.

[Core Function] Wanx 2.1 I2V Turbo is an older generation image-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2V.

[Core Function] Wanx 2.1 I2V Plus is an older generation image-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2V.

[Core Function] Wan 2.2 I2V Plus is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2V.

[Core Function] Wan 2.2 I2V Flash is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2V.

[Core Function] Wanx 2.5 I2V Preview is an older generation image-to-video model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2V.

[Core Function] Wan 2.6 I2V is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2V.

[Core Function] Wan 2.6 I2V Flash is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2V.