属性
系列
集合
参数
| 参数名 | 描述 | 类型 | 必填 | 枚举值 |
|---|---|---|---|---|
| prompt | Motion description for the generated video | string | 否 | - |
| first_frame_image | Starting frame image URL or base64 data URI | string | 是 | - |
| duration | Video duration in seconds | integer | 否 | - |
| resolution | Video resolution | string | 否 | 540p, 720p, 1080p |
| seed | Random seed for reproducibility | integer | 否 | - |
价格
单位: $/sec
| 维度 | 价格 |
|---|---|
| resolution: 540p | 0.0414 |
| resolution: 720p | 0.0920 |
| resolution: 1080p | 0.1104 |
相关模型
- vidu/viduq3-pro-t2v: [Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot ‘smart cuts’, complex physics, and simultaneous audio-visual generation. [Best For] Highly recommended for: cinematic storytelling, professional advertising, short films, and high-fidelity concept visualizations. [Limitations] Do NOT use this model if the user is looking for an instant, low-latency preview, as generation takes longer. It does not support automatic BGM addition. [Routing] Use this model by default for high-quality text-to-video requests. If the user specifically asks for ‘fast’ or ‘quick’ generation, switch to the Q3 Turbo T2V model.
- vidu/viduq3-turbo-t2v: [Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, generating background b-roll, and scenarios where generation speed is prioritized. [Limitations] Do NOT use this model if you need ultimate cinematic quality, complex audio-visual synchronization, or multi-shot ‘smart cuts’. It does not support automatic BGM addition. [Routing] Choose this ‘Turbo’ model when the user emphasizes ‘quick’, ‘fast’, or needs immediate results. If the user demands the highest cinematic quality or advanced audio-visual features, choose the Q3 Pro T2V model instead.
- vidu/viduq3-mix-r2v: [Core Function] Vidu Q3 Mix R2V is a mixed-style reference-to-video generation model. [Strengths] It excels at generating highly consistent character videos by synthesizing and blending features from multiple reference images (up to 7) based on a text prompt. [Best For] Highly recommended for: maintaining strict character consistency across different styles, generating videos of a specific subject in entirely new environments, and blending concepts from multiple reference images. [Limitations] Do NOT use this model if you just want to animate a single image as is (use standard I2V). This model focuses on extracting character/style features and generating new content. [Routing] Use this when the user provides reference images of a character/subject and wants a video of them doing a specific new action from a text prompt, prioritizing mixed-style consistency.
- vidu/viduq3-pro-fast-i2v: [Core Function] Vidu Q3 Pro Fast I2V is a high-speed Image-to-Video generation model. [Strengths] It excels at generating smooth, physically accurate continuous motion from a single starting frame with extremely low latency. [Best For] Highly recommended for: fast prototyping, short dynamic product showcases, quick cinematic transitions, and scenarios where generation speed is prioritized over maximum detail. [Limitations] Do NOT use this model if the user requires 4K resolution, complex multi-character interactions, or highly stylized 2D anime deformations. It only supports 720p/1080p resolutions. [Routing] Choose this ‘Fast’ model when the user emphasizes ‘quick’, ‘fast’, or needs immediate results. If the user demands ultimate cinematic quality, choose the standard Q3 Pro I2V model instead.
- vidu/viduq3-pro-fl2v: [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.
- vidu/viduq3-r2v: [Core Function] Vidu Q3 R2V is a high-quality reference-to-video generation model. [Strengths] It excels at generating detailed, cinematic videos that precisely follow a text prompt while highly preserving the character identity from provided reference images. [Best For] Highly recommended for: professional character-driven storytelling, high-fidelity avatar generation in new scenes, and cinematic films requiring consistent actors. [Limitations] Do NOT use this model if you just want to add motion to an existing image (use I2V). This model creates new scenes based on the prompt while keeping the character. [Routing] Use this by default when the user wants to generate a video of a specific character (provided via image) doing something new (provided via text prompt).
- vidu/viduq3-turbo-fl2v: [Core Function] Vidu Q3 Turbo FL2V is a fast First-Last frame transition video model. [Strengths] It excels at rapidly generating a smooth video transition bridging a specific starting image and an ending image. [Best For] Highly recommended for: quick visual morphs, before-and-after transitions, time-lapse simulations, and rapid storyboard filling. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. Do NOT use if you need the highest possible cinematic detail. [Routing] Use this when the user provides exactly two images and wants a fast transition between them. For higher quality transitions, use Q3 Pro FL2V.
- vidu/viduq3-turbo-i2v: [Core Function] Vidu Q3 Turbo I2V is a fast Image-to-Video generation model. [Strengths] It excels at quickly animating a starting frame into a video sequence with strong motion dynamics and low latency. [Best For] Highly recommended for: rapid social media content creation, quick animatics, and fast visual iterations from reference images. [Limitations] Do NOT use this model if you require the absolute highest cinematic fidelity or complex audio-visual synchronization. [Routing] Choose this model for general quick image-to-video tasks. For the highest quality, route to Q3 Pro I2V.
- vidu/viduq3-turbo-r2v: [Core Function] Vidu Q3 Turbo R2V is a fast reference-to-video generation model. [Strengths] It excels at quickly generating dynamic videos based on a text prompt while preserving the identity of the subjects from provided reference images. [Best For] Highly recommended for: rapid character animation, quick conceptual mockups with specific subjects, and fast social media content featuring consistent characters. [Limitations] Do NOT use this model if you need the absolute highest cinematic quality or if you just want to animate an image directly without a text prompt. [Routing] Choose this model for fast generation of a character performing actions based on a text prompt. For better quality, use Q3 R2V or Q3 Mix R2V.
- vidu/viduq3-drama: [Core Function] Vidu Q3 Drama (Short Play) is a script-to-video model that turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass. [Strengths] It excels at multi-shot narrative coherence, automatic storyboarding and cinematography, and keeping the identity of referenced characters, scenes, and props consistent across every shot. [Best For] Highly recommended for: scripted short dramas and web-series episodes, narrative short-form ads and brand stories, rapid storyboard and pre-visualization, and character-driven reels built from a cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image as-is (use a standard image-to-video model such as Vidu Q3 Pro instead); it requires a script and 1-14 reference assets, supports only 8-12 second clips at 1080p in 16:9 or 9:16, and is not intended for pixel-perfect single-product shots or complex multi-object physics. [Routing] Choose Vidu Q3 Drama when the user provides a script or narrative beats plus character/scene/prop references and wants an automatically directed multi-shot short play. If the user only wants to animate a single image or needs one continuous clip without scripted scene changes, route to a standard image-to-video model such as Vidu Q3 Pro.
- vidu/viduq3-ad: [Core Function] Vidu Q3 AD is a keyframe-driven short-play (short drama) Image-to-Video model that turns a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass. [Strengths] It excels at multi-shot narrative coherence and keeping referenced characters, scenes, and props consistent across shots, driven directly from keyframe reference images without manual shot-by-shot prompting. [Best For] Highly recommended for: short ad films and brand stories, scripted short dramas, multi-scene narrative clips generated from a script, and character-driven reels built from a small cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image (use a standard image-to-video model such as Vidu Q3 Pro/Turbo I2V instead); it requires script_content in traditional screenplay format (scene + characters + dialogue) plus 1-14 reference assets each with an image_uri, outputs 1080p in 16:9 or 9:16, and does not expose per-shot duration or style controls (duration is auto-planned). Asset type must be one of character/scene/tool. [Routing] Choose Vidu Q3 AD when the user provides a script and reference images and wants an automatically directed multi-shot short drama from keyframes. For animating a single image into one continuous clip, route to viduq3-pro-i2v / viduq3-turbo-i2v; for the standard short-play flow with placement/quality/duration/style controls, route to Vidu Q3 Drama.
- kling/kling-v3-i2v: [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.
- bytedance/seedance-2.0-i2v: [Core Function] Seedance 2.0 I2V is ByteDance’s flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.
- google/veo-3.1-i2v: [Core Function] Veo 3.1 I2V is Google’s cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animating concept art, creating cinematic transitions between images, and high-end video production. [Limitations] When using first+last frame or reference-only modes, the
durationparameter must strictly be 8. Negative prompts are not supported in reference-only mode. [Routing] Use this model by default for high-quality image-to-video tasks or when multiple reference images are provided. - minimax/hailuo-2.3-i2v: [Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.
- skywork/skyreels-i2v: [Core Function] SkyReels Image-to-Video animates one or more keyframe images into a video guided by a text prompt. [Strengths] Supports a start frame, an end frame, and tagged mid-frames for keyframe control; optional audio, 480p/720p/1080p output, and fast/std modes. [Best For] Bringing a photo to life, first-last-frame transitions, and keyframe-driven storyboards. [Limitations] Do NOT use this for pure text-to-video (use skyreels-t2v) or for editing an existing video (use the Omni / video-to-video models). It requires at least one of first_frame_image, end_frame_image, or mid_frame_images; output is capped at 1080p and 15s, and fast mode supports only sound=false. [Routing] Provide first_frame_image to animate from a start image, add end_frame_image for a transition, or supply mid_frame_images (each tag must appear in the prompt as @tag) for keyframe guidance.
- pixverse/motion-control: [Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limitations] Do NOT use this if you only have a video and no subject image (use Restyle or Extend instead), or if you need 1080p output (only 360p/540p/720p are supported). It requires BOTH a subject image (with a clear person or animal) AND a reference video (with a person as the primary focus). [Routing] Use when the user has one subject image and one motion reference video and wants the subject to mimic that motion.
- vidu/motion-sync: [Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target’s identity. [Best For] Highly recommended for: creating dance videos with custom characters, transferring complex choreography, and replicating specific physical actions onto avatars. [Limitations] Do NOT use this model if you want to change what a character is saying (use Lip Sync). It requires both a reference video for motion and a target image for appearance. [Routing] Use this model specifically when the user wants to copy the body movements or actions from one video onto a different character.
- google/gemini-omni-flash-t2v: [Core Function] Gemini Omni Flash T2V is Google’s fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.
- google/gemini-omni-flash-i2v: [Core Function] Gemini Omni Flash I2V is a fast Image-to-Video model that animates a single input image into a short 720p video via the Interactions API. [Strengths] It uses the provided image as the opening frame and generates smooth motion with natively synchronized audio at low latency. [Best For] Highly recommended for: bringing a still photo to life, quick product or portrait animation, and short social clips derived from a single image. [Limitations] Do NOT use this model if you need 1080p or 4K output, clips longer than 10 seconds, or the fusion of multiple reference images; it takes exactly one image and outputs 720p up to 10 seconds (16:9 or 9:16). Do NOT use it to edit an existing video. [Routing] Choose this when the user provides one image to animate. To fuse multiple reference images use Gemini Omni Flash R2V; to edit an existing video use Gemini Omni Flash Video Edit; for 4K cinematic results use Veo 3.1 I2V.
- google/gemini-omni-flash-r2v: [Core Function] Gemini Omni Flash R2V (Reference-to-Video) generates a short 720p video guided by up to three reference images via the Interactions API. [Strengths] It fuses the styles, subjects, or elements from multiple reference images (referred to in the text prompt) into a single coherent animated clip with synchronized audio. [Best For] Highly recommended for: blending characters or visual styles from several images, reference-guided creative shots, and multi-subject compositions where the prompt directs how the references combine. [Limitations] Do NOT use this model if you only have a single starting frame (use I2V instead), or if you need 1080p or 4K or clips longer than 10 seconds; it accepts 1 to 3 reference images and outputs 720p up to 10 seconds (16:9 or 9:16). [Routing] Choose this when the user supplies multiple reference images to combine into one video. For single first-frame animation use Gemini Omni Flash I2V; to modify an existing video use Gemini Omni Flash Video Edit.
- kling/kling-v3-omni-i2v: [Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.
- kling/kling-v3-turbo-i2v: [Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.








