What Vidu Q3 R2V Does
Vidu Q3 R2V is a high-quality reference-to-video model that generates new cinematic scenes from text while preserving the identity of one or more provided reference subjects. It is best for character-driven videos where the subject must stay recognizable while performing a new action in a new environment.
Key Features
- Preserves character identity from 1 to 7 reference images.
- Generates prompt-guided videos rather than simply animating the original image.
- Supports 3 to 16 second video duration.
- Offers 16:9, 9:16, and 1:1 aspect ratios.
- Supports 540p, 720p, and 1080p output with optional seed control.
How to Use Vidu Q3 R2V
- Prepare 1 to 7 clear portrait or subject reference images.
- Write a prompt describing the new scene, action, setting, camera direction, and mood.
- Submit a POST request to
/vidu/viduq3-r2vwithpromptandreference_images. - Optionally set
duration,aspect_ratio,resolution, andseed. - Poll the returned
get_resultURL until the asynchronous task completes.
Best Use Cases
- Character-consistent storytelling and short films.
- Avatar or actor reuse across new scenes.
- Brand mascot and spokesperson videos.
- Cinematic concept previews with a fixed subject.
- Multi-reference subject preservation for creative production.
Tips for Better Results
- Use sharp, front-facing reference images with clear subject features.
- Describe the target action and environment clearly instead of relying on vague style words.
- Choose R2V when you need identity preservation in a new scene, not simple image animation.
- Use a fixed seed when iterating on prompts.
- Start with 720p and shorter durations for testing, then increase quality for final outputs.
参数
| 参数名 | 描述 | 类型 | 必填 | 枚举值 |
|---|---|---|---|---|
| prompt | Video description text | string | 是 | - |
| reference_images | Portrait or subject images (1-7) whose appearance will be preserved in the generated video. Accepts URLs or base64 data URIs. | string[] | 是 | - |
| aspect_ratio | Video aspect ratio | string | 否 | 16:9, 9:16, 1:1 |
| duration | Video duration in seconds (minimum 3) | integer | 否 | - |
| resolution | Video resolution | string | 否 | 540p, 720p, 1080p |
| seed | Random seed for reproducibility | integer | 否 | - |
价格
单位: $/sec
| 维度 | 价格 |
|---|---|
| resolution: 540p | 0.0322 |
| resolution: 720p | 0.0552 |
| resolution: 1080p | 0.0690 |
相关模型
- vidu/viduq3-pro-t2v: [Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot ‘smart cuts’, complex physics, and simultaneous audio-visual generation. [Best For] Highly recommended for: cinematic storytelling, professional advertising, short films, and high-fidelity concept visualizations. [Limitations] Do NOT use this model if the user is looking for an instant, low-latency preview, as generation takes longer. It does not support automatic BGM addition. [Routing] Use this model by default for high-quality text-to-video requests. If the user specifically asks for ‘fast’ or ‘quick’ generation, switch to the Q3 Turbo T2V model.
- vidu/viduq3-turbo-t2v: [Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, generating background b-roll, and scenarios where generation speed is prioritized. [Limitations] Do NOT use this model if you need ultimate cinematic quality, complex audio-visual synchronization, or multi-shot ‘smart cuts’. It does not support automatic BGM addition. [Routing] Choose this ‘Turbo’ model when the user emphasizes ‘quick’, ‘fast’, or needs immediate results. If the user demands the highest cinematic quality or advanced audio-visual features, choose the Q3 Pro T2V model instead.
- vidu/viduq3-mix-r2v: [Core Function] Vidu Q3 Mix R2V is a mixed-style reference-to-video generation model. [Strengths] It excels at generating highly consistent character videos by synthesizing and blending features from multiple reference images (up to 7) based on a text prompt. [Best For] Highly recommended for: maintaining strict character consistency across different styles, generating videos of a specific subject in entirely new environments, and blending concepts from multiple reference images. [Limitations] Do NOT use this model if you just want to animate a single image as is (use standard I2V). This model focuses on extracting character/style features and generating new content. [Routing] Use this when the user provides reference images of a character/subject and wants a video of them doing a specific new action from a text prompt, prioritizing mixed-style consistency.
- vidu/viduq3-pro-fast-i2v: [Core Function] Vidu Q3 Pro Fast I2V is a high-speed Image-to-Video generation model. [Strengths] It excels at generating smooth, physically accurate continuous motion from a single starting frame with extremely low latency. [Best For] Highly recommended for: fast prototyping, short dynamic product showcases, quick cinematic transitions, and scenarios where generation speed is prioritized over maximum detail. [Limitations] Do NOT use this model if the user requires 4K resolution, complex multi-character interactions, or highly stylized 2D anime deformations. It only supports 720p/1080p resolutions. [Routing] Choose this ‘Fast’ model when the user emphasizes ‘quick’, ‘fast’, or needs immediate results. If the user demands ultimate cinematic quality, choose the standard Q3 Pro I2V model instead.
- vidu/viduq3-pro-fl2v: [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.
- vidu/viduq3-pro-i2v: [Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended for: bringing concept art to life, professional film production, high-end commercial showcases, and creating immersive environments from still images. [Limitations] Do NOT use this model if you need instant/real-time generation, as rendering takes longer. It does not support 4K resolution. [Routing] Use this model by default for high-quality image-to-video requests. If the user requires faster generation, route to Q3 Pro Fast or Q3 Turbo.
- vidu/viduq3-turbo-fl2v: [Core Function] Vidu Q3 Turbo FL2V is a fast First-Last frame transition video model. [Strengths] It excels at rapidly generating a smooth video transition bridging a specific starting image and an ending image. [Best For] Highly recommended for: quick visual morphs, before-and-after transitions, time-lapse simulations, and rapid storyboard filling. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. Do NOT use if you need the highest possible cinematic detail. [Routing] Use this when the user provides exactly two images and wants a fast transition between them. For higher quality transitions, use Q3 Pro FL2V.
- vidu/viduq3-turbo-i2v: [Core Function] Vidu Q3 Turbo I2V is a fast Image-to-Video generation model. [Strengths] It excels at quickly animating a starting frame into a video sequence with strong motion dynamics and low latency. [Best For] Highly recommended for: rapid social media content creation, quick animatics, and fast visual iterations from reference images. [Limitations] Do NOT use this model if you require the absolute highest cinematic fidelity or complex audio-visual synchronization. [Routing] Choose this model for general quick image-to-video tasks. For the highest quality, route to Q3 Pro I2V.
- vidu/viduq3-turbo-r2v: [Core Function] Vidu Q3 Turbo R2V is a fast reference-to-video generation model. [Strengths] It excels at quickly generating dynamic videos based on a text prompt while preserving the identity of the subjects from provided reference images. [Best For] Highly recommended for: rapid character animation, quick conceptual mockups with specific subjects, and fast social media content featuring consistent characters. [Limitations] Do NOT use this model if you need the absolute highest cinematic quality or if you just want to animate an image directly without a text prompt. [Routing] Choose this model for fast generation of a character performing actions based on a text prompt. For better quality, use Q3 R2V or Q3 Mix R2V.
- vidu/viduq3-drama: [Core Function] Vidu Q3 Drama (Short Play) is a script-to-video model that turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass. [Strengths] It excels at multi-shot narrative coherence, automatic storyboarding and cinematography, and keeping the identity of referenced characters, scenes, and props consistent across every shot. [Best For] Highly recommended for: scripted short dramas and web-series episodes, narrative short-form ads and brand stories, rapid storyboard and pre-visualization, and character-driven reels built from a cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image as-is (use a standard image-to-video model such as Vidu Q3 Pro instead); it requires a script and 1-14 reference assets, supports only 8-12 second clips at 1080p in 16:9 or 9:16, and is not intended for pixel-perfect single-product shots or complex multi-object physics. [Routing] Choose Vidu Q3 Drama when the user provides a script or narrative beats plus character/scene/prop references and wants an automatically directed multi-shot short play. If the user only wants to animate a single image or needs one continuous clip without scripted scene changes, route to a standard image-to-video model such as Vidu Q3 Pro.
- vidu/viduq3-ad: [Core Function] Vidu Q3 AD is a keyframe-driven short-play (short drama) Image-to-Video model that turns a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass. [Strengths] It excels at multi-shot narrative coherence and keeping referenced characters, scenes, and props consistent across shots, driven directly from keyframe reference images without manual shot-by-shot prompting. [Best For] Highly recommended for: short ad films and brand stories, scripted short dramas, multi-scene narrative clips generated from a script, and character-driven reels built from a small cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image (use a standard image-to-video model such as Vidu Q3 Pro/Turbo I2V instead); it requires script_content in traditional screenplay format (scene + characters + dialogue) plus 1-14 reference assets each with an image_uri, outputs 1080p in 16:9 or 9:16, and does not expose per-shot duration or style controls (duration is auto-planned). Asset type must be one of character/scene/tool. [Routing] Choose Vidu Q3 AD when the user provides a script and reference images and wants an automatically directed multi-shot short drama from keyframes. For animating a single image into one continuous clip, route to viduq3-pro-i2v / viduq3-turbo-i2v; for the standard short-play flow with placement/quality/duration/style controls, route to Vidu Q3 Drama.
- google/nano-banana-2: [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism (use the Imagen 4 series instead). [Routing] Route to this model for ‘fast’, ‘creative’, or ‘stylized’ high-volume requests.
- google/nano-banana-pro: [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [Limitations] Do NOT use for absolute photorealism (use Imagen 4). [Routing] Use this model for high-quality, creative, non-photorealistic requests.
- kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image generation model. [Strengths] It delivers state-of-the-art aesthetic quality, high-resolution outputs (up to 2K), and exceptional prompt adherence. [Best For] Highly recommended for: professional concept art, photorealistic portraits, and high-fidelity image generation. [Limitations] Do NOT use this model if you need to strictly reference or edit an existing image. [Routing] Use this as the default model for all text-to-image requests on the Kling platform.
- minimax/minimax-image-01-t2i: [Core Function] MiniMax Image-01 T2I is a multimodal text-to-image generation model. [Strengths] It excels at blending high-quality image generation with visual reasoning, allowing for strong prompt adherence and structural understanding. [Best For] Highly recommended for: general image generation, conceptual illustrations, and generating multiple images in a single batch. [Limitations] Do NOT use this model if the user specifically requests video generation. [Routing] Use this as the default text-to-image model for MiniMax API integrations.
- minimax/hailuo-2.3-i2v: [Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.
- minimax/hailuo-2.3-t2v: [Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.
- pixverse/video-restyle: [Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it only re-renders the visual style of an existing video. It requires an input video, and you must provide EITHER restyle_id (a preset style code from the PixVerse restyle list) OR restyle_prompt (free-text style, max 2048 chars), not both. [Routing] Use when the user wants to change the look of an existing video. Use restyle_id for an official preset, restyle_prompt for a custom style.




