Explore 29 production-ready Video to Video AI models on Modellix. Compare capabilities, try models in the playground, and integrate them through one unified API.

[Core Function] Gemini Omni Flash Video Edit performs conversational, instruction-driven editing of an existing video via the Interactions API. [Strengths] It applies natural-language edits (changing the scene, mood, style, lighting, background, or time of day) to an input video while preserving the source video's length and aspect ratio, with synchronized audio. [Best For] Highly recommended for: re-styling or re-lighting an existing clip, changing a video's setting or atmosphere, and quick instruction-based revisions of a short video. [Limitations] Do NOT use this model to generate a video from scratch (use T2V, I2V, or R2V), and do NOT expect to change the output resolution, aspect ratio, or duration: the output preserves the source video's aspect ratio and length, and the model does not accept aspectRatio or duration parameters. The source video should be 3 to 10 seconds. [Routing] Choose this only when the user provides an existing video to modify. To create a new video from text or images, use Gemini Omni Flash T2V, I2V, or R2V instead.

**[Core Function]** SkyReels Lip Sync (retalking) re-drives a talking video so the subject's lips match a given audio track. **[Strengths]** Accurate lip re-synchronization on an existing talking-head video. **[Best For]** Dubbing, re-voicing talking-head video, and localizing spoken video. **[Limitations]** Do NOT use this to generate new motion or content from scratch; it only re-times lips on an existing video. Requires video_url and audio_url. Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and the target audio_url; optionally provide reference_char_url to guide the driven face.

**[Core Function]** SkyReels Shot-Switching Extension continues a video while transitioning to a new shot or camera angle. **[Strengths]** Cinematic shot transitions (cut-in, cut-out, reverse-shot, multi-angle, cut-away) when extending footage. **[Best For]** Adding a new shot after existing footage and cinematic transitions. **[Limitations]** Do NOT use this for a plain single-shot continuation (use video-extension-single-shot) or for generation from scratch. Requires a prefix_video (mp4 URL); adds 2-5s. **[Routing]** Provide prompt and prefix_video; choose cut_type for the transition style, or Auto to let the model decide.

**[Core Function]** SkyReels Single-Shot Extension continues an existing single-shot video, generating additional seconds guided by a text prompt. **[Strengths]** Seamless single-shot continuation of the existing motion and scene. **[Best For]** Lengthening clips and continuing an action within one continuous shot. **[Limitations]** Do NOT use this to create a video from scratch (use skyreels-t2v) or to switch shots / add transitions (use video-extension-shot-switching). Requires a prefix_video (mp4 URL); adds 5-30s. **[Routing]** Provide prompt and prefix_video; set duration for how many seconds (5-30) to append.

**[Core Function]** SkyReels Restyle re-renders an existing video into a preset visual style. **[Strengths]** Consistent style transfer across all frames into a chosen named art style. **[Best For]** Turning footage into simpsons, lego, paper-cutting, amigurumi, animal-crossing, van-gogh, or pixel-art looks. **[Limitations]** Do NOT use this to change content, motion, or add new scenes; it only restyles an existing video (input <=30s). Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and a style_name from the supported list.

**[Core Function]** SkyReels Omni is a reference-driven video model that generates or edits video using reference images (@image) and/or a reference video (@video), bound by tags in the prompt. **[Strengths]** A single endpoint covers motion reference, subject/background replacement, object insertion/removal, local editing, and video extension. **[Best For]** Video subject or background swap, motion transfer onto an image, object add/remove, local video edits, and extending a reference video. **[Limitations]** Do NOT use this for pure text-to-video (use skyreels-t2v) or simple single-image animation (use skyreels-i2v). Each ref tag must appear in the prompt as @tag; ref_videos supports only one video (<=15s); a reference-type video may combine only with image-type ref_images, while an extend-type video cannot combine with ref_images. When ref_videos is provided, aspect_ratio is ignored (output matches the video). **[Routing]** Provide ref_images (type grid or image) for image references and/or a single ref_videos entry (type reference for motion/edit, type extend for continuation); the reference tags must be used in the prompt.

[Core Function] PixVerse Upscale increases the resolution and clarity of an existing video. [Strengths] Sharper detail and higher-resolution output without changing content. [Best For] Enhancing low-resolution footage, finalizing clips for delivery. [Limitations] Requires an input video; it enhances quality, it does NOT change content, style, or motion. [Routing] Use when the user wants to improve the resolution/quality of an existing video, not to generate, restyle, or extend content.

[Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limitations] Do NOT use this if you only have a video and no subject image (use Restyle or Extend instead), or if you need 1080p output (only 360p/540p/720p are supported). It requires BOTH a subject image (with a clear person or animal) AND a reference video (with a person as the primary focus). [Routing] Use when the user has one subject image and one motion reference video and wants the subject to mimic that motion.

[Core Function] PixVerse v6 Extend continues an existing video, generating additional seconds guided by a text prompt. [Strengths] Seamless continuation of the existing motion and scene. [Best For] Lengthening clips, continuing an action, adding an ending to footage. [Limitations] Do NOT use this to create a video from scratch (use Text-to-Video) or to restyle (use Restyle). It requires an input video; both duration and quality are required, with duration limited to 1-15 seconds added per call and output resolution up to 1080p. [Routing] Use when the user wants to make an existing video longer or continue its action.

[Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it only re-renders the visual style of an existing video. It requires an input video, and you must provide EITHER restyle_id (a preset style code from the PixVerse restyle list) OR restyle_prompt (free-text style, max 2048 chars), not both. [Routing] Use when the user wants to change the look of an existing video. Use restyle_id for an official preset, restyle_prompt for a custom style.

[Core Function] PixVerse Lip Sync drives a talking video so the subject's lips match given audio or text-to-speech. [Strengths] Accurate lip synchronization for talking-head videos; supports either an existing audio track or TTS from a chosen speaker voice. [Best For] Dubbing, virtual presenters, character dialogue, localizing spoken video. [Limitations] Do NOT use this to generate new motion or change content; it only re-times the subject's lips on an existing video. It requires an input video, and you must provide EITHER audio_url OR (speaker_id + tts_content), not both. [Routing] Use when the user has a video and wants the speaker's lips to match speech. Use audio_url for an existing voice track; use speaker_id (a named voice code) + tts_content (max 140 chars) to synthesize speech from text.

[Core Function] Seedance 2.0 Mini V2V is the lightweight video-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with required reference video and optional text/image/audio references. [Routing] Use for cost-efficient video-to-video transformations.

[Core Function] Grok Imagine Video Extend continues an existing video, generating additional footage beyond its end. [Strengths] It excels at seamlessly extending a clip with new prompt-guided motion. [Best For] Highly recommended for: lengthening short clips, continuing a scene, and adding follow-on action. [Limitations] Do NOT use the `duration` parameter expecting it to set the total video length; it only controls the length of the appended segment (2-10 seconds). Input video constraints are enforced by the upstream provider. [Routing] Use this model when the user wants to make a video longer. To restyle or modify an existing video, use Video Edit.

[Core Function] Grok Imagine Video Edit applies a prompt-guided transformation to an input video. [Strengths] It excels at restyling and modifying an existing video while keeping its original duration and aspect ratio. [Best For] Highly recommended for: restyling clips, applying visual effects, and prompt-driven video edits. [Limitations] Do NOT use this model to change the duration, aspect ratio, or resolution; the output preserves the input video's duration and aspect ratio, and those parameters are not configurable. Input video constraints (e.g. format/length) are enforced by the upstream provider. [Routing] Use this model when the user provides a video and wants it edited/restyled. To make a video longer, use Video Extend.

[Core Function] Vidu One-Click Trending Replicate is a viral video style cloning model. [Strengths] It excels at analyzing a trending or viral reference video and recreating its specific visual style, transitions, and pacing using the user's own provided subject images. [Best For] Highly recommended for: participating in social media video trends, quickly cloning viral visual effects, and applying popular editing styles to personal photos. [Limitations] Do NOT use this model for original cinematic storytelling. It is strictly meant to mimic the style of a provided reference video. [Routing] Use this when the user explicitly provides a 'trend' or 'viral' video and wants to recreate that exact vibe or transition style with their own images.

[Core Function] Vidu Lip Sync is a video-to-video audio synchronization model. [Strengths] It excels at reanimating lip movements in an existing video to precisely match a new replacement audio track, while preserving the original face identity. [Best For] Highly recommended for: dubbing videos into different languages, correcting spoken dialogue post-production, and creating realistic digital avatars. [Limitations] Do NOT use this model if you need to change body movements or generate a new video from scratch. It requires a pre-existing video and a clear audio track. [Routing] Use this model specifically when the user wants to change what a person in a video is saying. If the user wants to transfer body movements, use Vidu Motion Sync instead.

[Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target's identity. [Best For] Highly recommended for: creating dance videos with custom characters, transferring complex choreography, and replicating specific physical actions onto avatars. [Limitations] Do NOT use this model if you want to change what a character is saying (use Lip Sync). It requires both a reference video for motion and a target image for appearance. [Routing] Use this model specifically when the user wants to copy the body movements or actions from one video onto a different character.

[Core Function] HappyHorse 1.0 Video Edit is a streamlined video editing model. [Strengths] It provides high-quality video editing capabilities (with or without reference images) within the highly optimized HappyHorse architecture. [Best For] Highly recommended for: fast, high-quality video modifications, especially when the user explicitly requests HappyHorse. [Limitations] May lack the deep instruction-based logical replacement mechanics of Wan 2.7 Video Editing. [Routing] Route to this model when the user explicitly requests 'HappyHorse' for their video editing task.

[Core Function] Kling V3 Omni Video is the flagship unified multimodal video generation endpoint. [Strengths] It supports multi-shot narratives (up to 6 shots), 15-second durations, native audio, video editing (via refer_type 'base'), and character/element consistency across shots using `<image_N>` syntax. [Best For] Highly recommended for: professional video orchestration, multi-shot cinematic sequences, and maintaining strict character consistency. [Limitations] Do NOT use this model for simple, single-shot T2V/I2V tasks where the standard V3 model is more straightforward and cheaper. [Routing] Use this model whenever the user requests 'multi-shot', 'consistent characters', 'video editing', or complex multimodal inputs.

[Core Function] Kling Video O1 is the world's first reasoning-enhanced video model. [Strengths] It performs deep planning over the prompt before generation, delivering best-in-class physical consistency, complex motion logic, and strict adherence to long-form semantics. [Best For] Highly recommended for: complex physical interactions, logically demanding scenes, and prompts requiring deep reasoning. [Limitations] Do NOT use this model if you need multi-shot generation or 15-second durations (it is capped at 10s). [Routing] Route to this model when the prompt involves complex physics, logical sequences, or intricate physical interactions where standard models hallucinate.

[Core Function] Wan 2.2 Animate Mix - Character Replacement in Video is a legacy video editing and motion transfer model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 Video Editing.

[Core Function] Wan 2.2 Animate Move - Motion Transfer is a legacy video editing and motion transfer model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 Video Editing.

[Core Function] Wan 2.7 Video Editing is an instruction-based video modification model. [Strengths] It supports complex video editing tasks like content replacement using reference images, and replicating actions, effects, and camera movements. [Best For] Highly recommended for: modifying existing video footage, style transfer on videos, and targeted element replacement. [Limitations] Do NOT use this model to generate a brand new video from scratch; it requires an input video. [Routing] Use this model by default whenever a user wants to edit, alter, or restyle an existing video.

[Core Function] Wan 2.1 VACE Plus - Unified Video Editing Model is an older generation image-to-image editing model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2I.

[Core Function] Wan 2.6 Reference-to-Video Flash is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7.

[Core Function] Wan 2.6 Reference-to-Video is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7.

[Core Function] Wan 2.7 Reference-to-Video is a highly capable character/entity reference video model. [Strengths] It natively supports entity reference, voice customization, and playbook-based video generation from a single storyboard. [Best For] Highly recommended for: creating consistent video series, brand mascot animation, and storyboard-driven storytelling. [Limitations] Do NOT use this model for simple, single-image direct animation (use I2V instead). [Routing] Use this by default for complex character consistency and storyboard generation tasks on Alibaba.

[Core Function] Seedance 2.0 Fast V2V is a high-speed video-to-video generation model. [Strengths] It offers rapid video transformation capabilities based on the 2.0 architecture. [Best For] Highly recommended for: quick video restyling and fast iterations. [Limitations] Do NOT use this model when absolute maximum visual quality is required (use standard 2.0). [Routing] Use this when speed is the primary concern for video transformations.

[Core Function] Seedance 2.0 V2V is ByteDance's flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] Highly recommended for: complex video-to-video transformations, restyling existing footage, and creating dynamic multi-shot edits. [Limitations] Do NOT use this model for simple still-image generation (use Seedream instead). [Routing] Use this model by default for any video editing or video-to-video generation tasks.