分类

2026 年最佳视频生视频 AI 模型

29 个模型更新于 Jul 2026

关于 视频生视频 模型

在 Modellix 探索 29 个生产可用的视频生视频 AI 模型,对比模型能力、在线试用,并通过统一 API 快速完成集成。

全部 视频生视频 模型

google/gemini-omni-flash-video-edit

google/gemini-omni-flash-video-edit

video-to-video

[Core Function] Gemini Omni Flash Video Edit performs conversational, instruction-driven editing of an existing video via the Interactions API. [Strengths] It applies natural-language edits (changing the scene, mood, style, lighting, background, or time of day) to an input video while preserving the source video's length and aspect ratio, with synchronized audio. [Best For] Highly recommended for: re-styling or re-lighting an existing clip, changing a video's setting or atmosphere, and quick instruction-based revisions of a short video. [Limitations] Do NOT use this model to generate a video from scratch (use T2V, I2V, or R2V), and do NOT expect to change the output resolution, aspect ratio, or duration: the output preserves the source video's aspect ratio and length, and the model does not accept aspectRatio or duration parameters. The source video should be 3 to 10 seconds. [Routing] Choose this only when the user provides an existing video to modify. To create a new video from text or images, use Gemini Omni Flash T2V, I2V, or R2V instead.

skywork/sky-lipsync

skywork/sky-lipsync

video-to-video

**[Core Function]** SkyReels Lip Sync (retalking) re-drives a talking video so the subject's lips match a given audio track. **[Strengths]** Accurate lip re-synchronization on an existing talking-head video. **[Best For]** Dubbing, re-voicing talking-head video, and localizing spoken video. **[Limitations]** Do NOT use this to generate new motion or content from scratch; it only re-times lips on an existing video. Requires video_url and audio_url. Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and the target audio_url; optionally provide reference_char_url to guide the driven face.

skywork/video-extension-shot-switching

skywork/video-extension-shot-switching

video-to-video

**[Core Function]** SkyReels Shot-Switching Extension continues a video while transitioning to a new shot or camera angle. **[Strengths]** Cinematic shot transitions (cut-in, cut-out, reverse-shot, multi-angle, cut-away) when extending footage. **[Best For]** Adding a new shot after existing footage and cinematic transitions. **[Limitations]** Do NOT use this for a plain single-shot continuation (use video-extension-single-shot) or for generation from scratch. Requires a prefix_video (mp4 URL); adds 2-5s. **[Routing]** Provide prompt and prefix_video; choose cut_type for the transition style, or Auto to let the model decide.

skywork/video-extension-single-shot

skywork/video-extension-single-shot

video-to-video

**[Core Function]** SkyReels Single-Shot Extension continues an existing single-shot video, generating additional seconds guided by a text prompt. **[Strengths]** Seamless single-shot continuation of the existing motion and scene. **[Best For]** Lengthening clips and continuing an action within one continuous shot. **[Limitations]** Do NOT use this to create a video from scratch (use skyreels-t2v) or to switch shots / add transitions (use video-extension-shot-switching). Requires a prefix_video (mp4 URL); adds 5-30s. **[Routing]** Provide prompt and prefix_video; set duration for how many seconds (5-30) to append.

skywork/video-restyling

skywork/video-restyling

video-to-video

**[Core Function]** SkyReels Restyle re-renders an existing video into a preset visual style. **[Strengths]** Consistent style transfer across all frames into a chosen named art style. **[Best For]** Turning footage into simpsons, lego, paper-cutting, amigurumi, animal-crossing, van-gogh, or pixel-art looks. **[Limitations]** Do NOT use this to change content, motion, or add new scenes; it only restyles an existing video (input <=30s). Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and a style_name from the supported list.

skywork/skyreels-omni

skywork/skyreels-omni

video-to-video

**[Core Function]** SkyReels Omni is a reference-driven video model that generates or edits video using reference images (@image) and/or a reference video (@video), bound by tags in the prompt. **[Strengths]** A single endpoint covers motion reference, subject/background replacement, object insertion/removal, local editing, and video extension. **[Best For]** Video subject or background swap, motion transfer onto an image, object add/remove, local video edits, and extending a reference video. **[Limitations]** Do NOT use this for pure text-to-video (use skyreels-t2v) or simple single-image animation (use skyreels-i2v). Each ref tag must appear in the prompt as @tag; ref_videos supports only one video (<=15s); a reference-type video may combine only with image-type ref_images, while an extend-type video cannot combine with ref_images. When ref_videos is provided, aspect_ratio is ignored (output matches the video). **[Routing]** Provide ref_images (type grid or image) for image references and/or a single ref_videos entry (type reference for motion/edit, type extend for continuation); the reference tags must be used in the prompt.

pixverse/upscale-video

pixverse/upscale-video

video-to-video

[Core Function] PixVerse Upscale increases the resolution and clarity of an existing video. [Strengths] Sharper detail and higher-resolution output without changing content. [Best For] Enhancing low-resolution footage, finalizing clips for delivery. [Limitations] Requires an input video; it enhances quality, it does NOT change content, style, or motion. [Routing] Use when the user wants to improve the resolution/quality of an existing video, not to generate, restyle, or extend content.

pixverse/motion-control

pixverse/motion-control

video-to-video

[Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limitations] Do NOT use this if you only have a video and no subject image (use Restyle or Extend instead), or if you need 1080p output (only 360p/540p/720p are supported). It requires BOTH a subject image (with a clear person or animal) AND a reference video (with a person as the primary focus). [Routing] Use when the user has one subject image and one motion reference video and wants the subject to mimic that motion.

pixverse/v6-video-extend

pixverse/v6-video-extend

video-to-video

[Core Function] PixVerse v6 Extend continues an existing video, generating additional seconds guided by a text prompt. [Strengths] Seamless continuation of the existing motion and scene. [Best For] Lengthening clips, continuing an action, adding an ending to footage. [Limitations] Do NOT use this to create a video from scratch (use Text-to-Video) or to restyle (use Restyle). It requires an input video; both duration and quality are required, with duration limited to 1-15 seconds added per call and output resolution up to 1080p. [Routing] Use when the user wants to make an existing video longer or continue its action.

pixverse/video-restyle

pixverse/video-restyle

video-to-video

[Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it only re-renders the visual style of an existing video. It requires an input video, and you must provide EITHER restyle_id (a preset style code from the PixVerse restyle list) OR restyle_prompt (free-text style, max 2048 chars), not both. [Routing] Use when the user wants to change the look of an existing video. Use restyle_id for an official preset, restyle_prompt for a custom style.

pixverse/lipsync

pixverse/lipsync

video-to-video

[Core Function] PixVerse Lip Sync drives a talking video so the subject's lips match given audio or text-to-speech. [Strengths] Accurate lip synchronization for talking-head videos; supports either an existing audio track or TTS from a chosen speaker voice. [Best For] Dubbing, virtual presenters, character dialogue, localizing spoken video. [Limitations] Do NOT use this to generate new motion or change content; it only re-times the subject's lips on an existing video. It requires an input video, and you must provide EITHER audio_url OR (speaker_id + tts_content), not both. [Routing] Use when the user has a video and wants the speaker's lips to match speech. Use audio_url for an existing voice track; use speaker_id (a named voice code) + tts_content (max 140 chars) to synthesize speech from text.

bytedance/seedance-2.0-mini-v2v

bytedance/seedance-2.0-mini-v2v

video-to-video

[Core Function] Seedance 2.0 Mini V2V is the lightweight video-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with required reference video and optional text/image/audio references. [Routing] Use for cost-efficient video-to-video transformations.

xai/grok-imagine-video-extend

xai/grok-imagine-video-extend

video-to-video

[Core Function] Grok Imagine Video Extend continues an existing video, generating additional footage beyond its end. [Strengths] It excels at seamlessly extending a clip with new prompt-guided motion. [Best For] Highly recommended for: lengthening short clips, continuing a scene, and adding follow-on action. [Limitations] Do NOT use the `duration` parameter expecting it to set the total video length; it only controls the length of the appended segment (2-10 seconds). Input video constraints are enforced by the upstream provider. [Routing] Use this model when the user wants to make a video longer. To restyle or modify an existing video, use Video Edit.

xai/grok-imagine-video-edit

xai/grok-imagine-video-edit

video-to-video

[Core Function] Grok Imagine Video Edit applies a prompt-guided transformation to an input video. [Strengths] It excels at restyling and modifying an existing video while keeping its original duration and aspect ratio. [Best For] Highly recommended for: restyling clips, applying visual effects, and prompt-driven video edits. [Limitations] Do NOT use this model to change the duration, aspect ratio, or resolution; the output preserves the input video's duration and aspect ratio, and those parameters are not configurable. Input video constraints (e.g. format/length) are enforced by the upstream provider. [Routing] Use this model when the user provides a video and wants it edited/restyled. To make a video longer, use Video Extend.

vidu/one-click-trending-replicate

vidu/one-click-trending-replicate

video-to-video

[Core Function] Vidu One-Click Trending Replicate is a viral video style cloning model. [Strengths] It excels at analyzing a trending or viral reference video and recreating its specific visual style, transitions, and pacing using the user's own provided subject images. [Best For] Highly recommended for: participating in social media video trends, quickly cloning viral visual effects, and applying popular editing styles to personal photos. [Limitations] Do NOT use this model for original cinematic storytelling. It is strictly meant to mimic the style of a provided reference video. [Routing] Use this when the user explicitly provides a 'trend' or 'viral' video and wants to recreate that exact vibe or transition style with their own images.

vidu/lip-sync

vidu/lip-sync

video-to-video

[Core Function] Vidu Lip Sync is a video-to-video audio synchronization model. [Strengths] It excels at reanimating lip movements in an existing video to precisely match a new replacement audio track, while preserving the original face identity. [Best For] Highly recommended for: dubbing videos into different languages, correcting spoken dialogue post-production, and creating realistic digital avatars. [Limitations] Do NOT use this model if you need to change body movements or generate a new video from scratch. It requires a pre-existing video and a clear audio track. [Routing] Use this model specifically when the user wants to change what a person in a video is saying. If the user wants to transfer body movements, use Vidu Motion Sync instead.

vidu/motion-sync

vidu/motion-sync

video-to-video

[Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target's identity. [Best For] Highly recommended for: creating dance videos with custom characters, transferring complex choreography, and replicating specific physical actions onto avatars. [Limitations] Do NOT use this model if you want to change what a character is saying (use Lip Sync). It requires both a reference video for motion and a target image for appearance. [Routing] Use this model specifically when the user wants to copy the body movements or actions from one video onto a different character.

alibaba/happyhorse-1.0-video-edit

alibaba/happyhorse-1.0-video-edit

video-to-video

[Core Function] HappyHorse 1.0 Video Edit is a streamlined video editing model. [Strengths] It provides high-quality video editing capabilities (with or without reference images) within the highly optimized HappyHorse architecture. [Best For] Highly recommended for: fast, high-quality video modifications, especially when the user explicitly requests HappyHorse. [Limitations] May lack the deep instruction-based logical replacement mechanics of Wan 2.7 Video Editing. [Routing] Route to this model when the user explicitly requests 'HappyHorse' for their video editing task.

kling/kling-v3-omni-video

kling/kling-v3-omni-video

video-to-video

[Core Function] Kling V3 Omni Video is the flagship unified multimodal video generation endpoint. [Strengths] It supports multi-shot narratives (up to 6 shots), 15-second durations, native audio, video editing (via refer_type 'base'), and character/element consistency across shots using `<image_N>` syntax. [Best For] Highly recommended for: professional video orchestration, multi-shot cinematic sequences, and maintaining strict character consistency. [Limitations] Do NOT use this model for simple, single-shot T2V/I2V tasks where the standard V3 model is more straightforward and cheaper. [Routing] Use this model whenever the user requests 'multi-shot', 'consistent characters', 'video editing', or complex multimodal inputs.

kling/kling-video-o1

kling/kling-video-o1

video-to-video

[Core Function] Kling Video O1 is the world's first reasoning-enhanced video model. [Strengths] It performs deep planning over the prompt before generation, delivering best-in-class physical consistency, complex motion logic, and strict adherence to long-form semantics. [Best For] Highly recommended for: complex physical interactions, logically demanding scenes, and prompts requiring deep reasoning. [Limitations] Do NOT use this model if you need multi-shot generation or 15-second durations (it is capped at 10s). [Routing] Route to this model when the prompt involves complex physics, logical sequences, or intricate physical interactions where standard models hallucinate.

alibaba/wan2.2-animate-mix

alibaba/wan2.2-animate-mix

video-to-video

[Core Function] Wan 2.2 Animate Mix - Character Replacement in Video is a legacy video editing and motion transfer model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 Video Editing.

alibaba/wan2.2-animate-move

alibaba/wan2.2-animate-move

video-to-video

[Core Function] Wan 2.2 Animate Move - Motion Transfer is a legacy video editing and motion transfer model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 Video Editing.

alibaba/wan2.7-videoedit

alibaba/wan2.7-videoedit

video-to-video

[Core Function] Wan 2.7 Video Editing is an instruction-based video modification model. [Strengths] It supports complex video editing tasks like content replacement using reference images, and replicating actions, effects, and camera movements. [Best For] Highly recommended for: modifying existing video footage, style transfer on videos, and targeted element replacement. [Limitations] Do NOT use this model to generate a brand new video from scratch; it requires an input video. [Routing] Use this model by default whenever a user wants to edit, alter, or restyle an existing video.

alibaba/wan2.1-vace-plus

alibaba/wan2.1-vace-plus

video-to-video

[Core Function] Wan 2.1 VACE Plus - Unified Video Editing Model is an older generation image-to-image editing model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 I2I.

alibaba/wan2.6-r2v-flash

alibaba/wan2.6-r2v-flash

video-to-video

[Core Function] Wan 2.6 Reference-to-Video Flash is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7.

alibaba/wan2.6-r2v

alibaba/wan2.6-r2v

video-to-video

[Core Function] Wan 2.6 Reference-to-Video is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7.

alibaba/wan2.7-r2v

alibaba/wan2.7-r2v

video-to-video

[Core Function] Wan 2.7 Reference-to-Video is a highly capable character/entity reference video model. [Strengths] It natively supports entity reference, voice customization, and playbook-based video generation from a single storyboard. [Best For] Highly recommended for: creating consistent video series, brand mascot animation, and storyboard-driven storytelling. [Limitations] Do NOT use this model for simple, single-image direct animation (use I2V instead). [Routing] Use this by default for complex character consistency and storyboard generation tasks on Alibaba.

bytedance/seedance-2.0-fast-v2v

bytedance/seedance-2.0-fast-v2v

video-to-video

[Core Function] Seedance 2.0 Fast V2V is a high-speed video-to-video generation model. [Strengths] It offers rapid video transformation capabilities based on the 2.0 architecture. [Best For] Highly recommended for: quick video restyling and fast iterations. [Limitations] Do NOT use this model when absolute maximum visual quality is required (use standard 2.0). [Routing] Use this when speed is the primary concern for video transformations.

bytedance/seedance-2.0-v2v

bytedance/seedance-2.0-v2v

video-to-video

[Core Function] Seedance 2.0 V2V is ByteDance's flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] Highly recommended for: complex video-to-video transformations, restyling existing footage, and creating dynamic multi-shot edits. [Limitations] Do NOT use this model for simple still-image generation (use Seedream instead). [Routing] Use this model by default for any video editing or video-to-video generation tasks.

没有找到需要的模型? 告诉我们。

探索更多

Gemini Omni
Gemini Omni
系列4 个模型
图生视频
分类77 个模型
文生视频
分类38 个模型
AI Animation Generator
AI Animation Generator
集合12 个模型
AI Anime Generator
AI Anime Generator
集合14 个模型
SkyReels
SkyReels
系列4 个模型