
[Core Function] Gemini Omni 1.1 Flash V2V edits an existing video from a text instruction. [Strengths] It can change scene, mood, style, lighting, or time of day while keeping the source length and aspect ratio, with optional 360p to 4k output and synchronized audio. [Best For] Highly recommended for: re-styling or re-lighting a clip, changing setting or atmosphere, and quick revisions of a short video. [Limitations] Do NOT use this model to generate a video from scratch (use Gemini Omni 1.1 Flash T2V or I2V). The source video should be 3 to 10 seconds; output length and aspect ratio follow the source. [Routing] Choose this only when the user provides an existing video to modify.

[Core Function] Gemini Omni 1.1 Flash I2V animates a single image into a short video with synchronized audio. [Strengths] It uses the input image as the opening frame and supports 360p, 720p, 1080p, and 4k output, 16:9 or 9:16, and 3 to 10 second clips. [Best For] Highly recommended for: animating a still, product or character motion from one frame, and social clips that need built-in audio at 1080p or 4k. [Limitations] Do NOT use this model if you have no starting image (use Gemini Omni 1.1 Flash T2V) or want to edit an existing video (use Gemini Omni 1.1 Flash V2V). Clips are capped at 10 seconds. [Routing] Choose this when the user provides one starting frame. For text-only generation, use Gemini Omni 1.1 Flash T2V.

[Core Function] Gemini Omni 1.1 Flash T2V generates a short video with synchronized audio from a text prompt. [Strengths] It supports 360p, 720p, 1080p, and 4k output, 16:9 or 9:16, and 3 to 10 second clips with native speech, music, and sound effects. [Best For] Highly recommended for: rapid prototyping, short social and marketing clips, concept visualization, and cases that need built-in audio at 1080p or 4k. [Limitations] Do NOT use this model if you need clips longer than 10 seconds, a starting image as the first frame (use Gemini Omni 1.1 Flash I2V), or edits to an existing video (use Gemini Omni 1.1 Flash V2V). [Routing] Choose this when the user wants a new video from text only. If they provide a starting image, use Gemini Omni 1.1 Flash I2V. For maximum cinematic control, choose Veo 3.1 T2V.

[Core Function] MiniMax H3 V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: motion remix, style transfer from video clips, keeping subject motion while changing the scene description, and multi-clip reference guidance. [Limitations] Do NOT use this model for text-only generation or first or last frame transitions. prompt, duration, resolution, and reference_videos are required. Do NOT send first_frame_image or last_frame_image on this endpoint. [Routing] Choose MiniMax H3 V2V when the user provides reference videos. For reference stills only, use MiniMax H3 I2V. For start or end frames, use MiniMax H3 FL2V. For text only, use MiniMax H3 T2V.

[Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions or when the primary input is a reference video. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 I2V when the user provides reference stills. For start or end frames, use MiniMax H3 FL2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

[Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 I2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

[Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.

[Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 480p/720p/1080p, 4-30 second output, and mp4 or mov containers. [Best For] Highly recommended for: editing existing clips, extending motion from a base video, multi-reference restyling, and prompt-driven composition that cites @Video1 or @Image1. [Limitations] Do NOT use this model for text-only or first-frame-only workflows. video_urls is required. Do NOT send first_frame_image or last_frame_image. Do NOT use this model when the user requires 4k output. Do NOT use this model for audio-only input. [Routing] Prefer this model for Seedance 2.5 edit, extension, and video-reference jobs. Use Seedance 2.5 T2V for text-only and Seedance 2.5 I2V for image-first generation.

[Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 480p/720p/1080p, 4-30 second duration, and mp4 or mov output. [Best For] Highly recommended for: animating a keyframe, first-to-last transitions, multi-image character consistency, and image-led storytelling on Seedance 2.5. [Limitations] Do NOT mix first_frame_image or last_frame_image with reference_images. Do NOT send video_urls on I2V. Do NOT use audio_urls alone. Do NOT use this model when the user requires 4k output. [Routing] Prefer this model for Seedance 2.5 image-driven generation. Use Seedance 2.5 T2V for text-only requests and Seedance 2.5 V2V when reference video is required.