vidu/viduq3-r2v

viduq3-r2v
文档
Schema

[Core Function] Vidu Q3 R2V is a high-quality reference-to-video generation model. [Strengths] It excels at generating detailed, cinematic videos that precisely follow a text prompt while highly preserving the character identity from provided reference images. [Best For] Highly recommended for: professional character-driven storytelling, high-fidelity avatar generation in new scenes, and cinematic films requiring consistent actors. [Limitations] Do NOT use this model if you just want to add motion to an existing image (use I2V). This model creates new scenes based on the prompt while keeping the character. [Routing] Use this by default when the user wants to generate a video of a specific character (provided via image) doing something new (provided via text prompt).

$0.0322~$0.0690/sec
image-to-video

输入

Video description text
Portrait or subject images (1-7) whose appearance will be preserved in the generated video. Accepts URLs or base64 data URIs.
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Video aspect ratio
16:9
Video duration in seconds (minimum 3)
Video resolution
720p
Random seed for reproducibility

结果

暂无结果

运行模型后,结果将在这里显示。

Next:

README

What Vidu Q3 R2V Does

Vidu Q3 R2V is a high-quality reference-to-video model that generates new cinematic scenes from text while preserving the identity of one or more provided reference subjects. It is best for character-driven videos where the subject must stay recognizable while performing a new action in a new environment.

Key Features

  • Preserves character identity from 1 to 7 reference images.
  • Generates prompt-guided videos rather than simply animating the original image.
  • Supports 3 to 16 second video duration.
  • Offers 16:9, 9:16, and 1:1 aspect ratios.
  • Supports 540p, 720p, and 1080p output with optional seed control.

How to Use Vidu Q3 R2V

  1. Prepare 1 to 7 clear portrait or subject reference images.
  2. Write a prompt describing the new scene, action, setting, camera direction, and mood.
  3. Submit a POST request to /vidu/viduq3-r2v with prompt and reference_images.
  4. Optionally set duration, aspect_ratio, resolution, and seed.
  5. Poll the returned get_result URL until the asynchronous task completes.

Best Use Cases

  • Character-consistent storytelling and short films.
  • Avatar or actor reuse across new scenes.
  • Brand mascot and spokesperson videos.
  • Cinematic concept previews with a fixed subject.
  • Multi-reference subject preservation for creative production.

Tips for Better Results

  • Use sharp, front-facing reference images with clear subject features.
  • Describe the target action and environment clearly instead of relying on vague style words.
  • Choose R2V when you need identity preservation in a new scene, not simple image animation.
  • Use a fixed seed when iterating on prompts.
  • Start with 720p and shorter durations for testing, then increase quality for final outputs.

参数

参数名 描述 类型 必填 枚举值
prompt Video description text string -
reference_images Portrait or subject images (1-7) whose appearance will be preserved in the generated video. Accepts URLs or base64 data URIs. string[] -
aspect_ratio Video aspect ratio string 16:9, 9:16, 1:1
duration Video duration in seconds (minimum 3) integer -
resolution Video resolution string 540p, 720p, 1080p
seed Random seed for reproducibility integer -

价格

单位: $/sec

维度 价格
resolution: 540p 0.0322
resolution: 720p 0.0552
resolution: 1080p 0.0690

相关模型

  • vidu/viduq3-pro-t2v: [Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot ‘smart cuts’, complex physics, and simultaneous audio-visual generation. [Best For] Highly recommended for: cinematic storytelling, professional advertising, short films, and high-fidelity concept visualizations. [Limitations] Do NOT use this model if the user is looking for an instant, low-latency preview, as generation takes longer. It does not support automatic BGM addition. [Routing] Use this model by default for high-quality text-to-video requests. If the user specifically asks for ‘fast’ or ‘quick’ generation, switch to the Q3 Turbo T2V model.
  • vidu/viduq3-turbo-t2v: [Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, generating background b-roll, and scenarios where generation speed is prioritized. [Limitations] Do NOT use this model if you need ultimate cinematic quality, complex audio-visual synchronization, or multi-shot ‘smart cuts’. It does not support automatic BGM addition. [Routing] Choose this ‘Turbo’ model when the user emphasizes ‘quick’, ‘fast’, or needs immediate results. If the user demands the highest cinematic quality or advanced audio-visual features, choose the Q3 Pro T2V model instead.
  • vidu/viduq3-mix-r2v: [Core Function] Vidu Q3 Mix R2V is a mixed-style reference-to-video generation model. [Strengths] It excels at generating highly consistent character videos by synthesizing and blending features from multiple reference images (up to 7) based on a text prompt. [Best For] Highly recommended for: maintaining strict character consistency across different styles, generating videos of a specific subject in entirely new environments, and blending concepts from multiple reference images. [Limitations] Do NOT use this model if you just want to animate a single image as is (use standard I2V). This model focuses on extracting character/style features and generating new content. [Routing] Use this when the user provides reference images of a character/subject and wants a video of them doing a specific new action from a text prompt, prioritizing mixed-style consistency.
  • vidu/viduq3-pro-fast-i2v: [Core Function] Vidu Q3 Pro Fast I2V is a high-speed Image-to-Video generation model. [Strengths] It excels at generating smooth, physically accurate continuous motion from a single starting frame with extremely low latency. [Best For] Highly recommended for: fast prototyping, short dynamic product showcases, quick cinematic transitions, and scenarios where generation speed is prioritized over maximum detail. [Limitations] Do NOT use this model if the user requires 4K resolution, complex multi-character interactions, or highly stylized 2D anime deformations. It only supports 720p/1080p resolutions. [Routing] Choose this ‘Fast’ model when the user emphasizes ‘quick’, ‘fast’, or needs immediate results. If the user demands ultimate cinematic quality, choose the standard Q3 Pro I2V model instead.
  • vidu/viduq3-pro-fl2v: [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.
  • vidu/viduq3-pro-i2v: [Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended for: bringing concept art to life, professional film production, high-end commercial showcases, and creating immersive environments from still images. [Limitations] Do NOT use this model if you need instant/real-time generation, as rendering takes longer. It does not support 4K resolution. [Routing] Use this model by default for high-quality image-to-video requests. If the user requires faster generation, route to Q3 Pro Fast or Q3 Turbo.
  • vidu/viduq3-turbo-fl2v: [Core Function] Vidu Q3 Turbo FL2V is a fast First-Last frame transition video model. [Strengths] It excels at rapidly generating a smooth video transition bridging a specific starting image and an ending image. [Best For] Highly recommended for: quick visual morphs, before-and-after transitions, time-lapse simulations, and rapid storyboard filling. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. Do NOT use if you need the highest possible cinematic detail. [Routing] Use this when the user provides exactly two images and wants a fast transition between them. For higher quality transitions, use Q3 Pro FL2V.
  • vidu/viduq3-turbo-i2v: [Core Function] Vidu Q3 Turbo I2V is a fast Image-to-Video generation model. [Strengths] It excels at quickly animating a starting frame into a video sequence with strong motion dynamics and low latency. [Best For] Highly recommended for: rapid social media content creation, quick animatics, and fast visual iterations from reference images. [Limitations] Do NOT use this model if you require the absolute highest cinematic fidelity or complex audio-visual synchronization. [Routing] Choose this model for general quick image-to-video tasks. For the highest quality, route to Q3 Pro I2V.
  • vidu/viduq3-turbo-r2v: [Core Function] Vidu Q3 Turbo R2V is a fast reference-to-video generation model. [Strengths] It excels at quickly generating dynamic videos based on a text prompt while preserving the identity of the subjects from provided reference images. [Best For] Highly recommended for: rapid character animation, quick conceptual mockups with specific subjects, and fast social media content featuring consistent characters. [Limitations] Do NOT use this model if you need the absolute highest cinematic quality or if you just want to animate an image directly without a text prompt. [Routing] Choose this model for fast generation of a character performing actions based on a text prompt. For better quality, use Q3 R2V or Q3 Mix R2V.
  • vidu/viduq3-drama: [Core Function] Vidu Q3 Drama (Short Play) is a script-to-video model that turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass. [Strengths] It excels at multi-shot narrative coherence, automatic storyboarding and cinematography, and keeping the identity of referenced characters, scenes, and props consistent across every shot. [Best For] Highly recommended for: scripted short dramas and web-series episodes, narrative short-form ads and brand stories, rapid storyboard and pre-visualization, and character-driven reels built from a cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image as-is (use a standard image-to-video model such as Vidu Q3 Pro instead); it requires a script and 1-14 reference assets, supports only 8-12 second clips at 1080p in 16:9 or 9:16, and is not intended for pixel-perfect single-product shots or complex multi-object physics. [Routing] Choose Vidu Q3 Drama when the user provides a script or narrative beats plus character/scene/prop references and wants an automatically directed multi-shot short play. If the user only wants to animate a single image or needs one continuous clip without scripted scene changes, route to a standard image-to-video model such as Vidu Q3 Pro.
  • vidu/viduq3-ad: [Core Function] Vidu Q3 AD is a keyframe-driven short-play (short drama) Image-to-Video model that turns a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass. [Strengths] It excels at multi-shot narrative coherence and keeping referenced characters, scenes, and props consistent across shots, driven directly from keyframe reference images without manual shot-by-shot prompting. [Best For] Highly recommended for: short ad films and brand stories, scripted short dramas, multi-scene narrative clips generated from a script, and character-driven reels built from a small cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image (use a standard image-to-video model such as Vidu Q3 Pro/Turbo I2V instead); it requires script_content in traditional screenplay format (scene + characters + dialogue) plus 1-14 reference assets each with an image_uri, outputs 1080p in 16:9 or 9:16, and does not expose per-shot duration or style controls (duration is auto-planned). Asset type must be one of character/scene/tool. [Routing] Choose Vidu Q3 AD when the user provides a script and reference images and wants an automatically directed multi-shot short drama from keyframes. For animating a single image into one continuous clip, route to viduq3-pro-i2v / viduq3-turbo-i2v; for the standard short-play flow with placement/quality/duration/style controls, route to Vidu Q3 Drama.
  • google/nano-banana-2: [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism (use the Imagen 4 series instead). [Routing] Route to this model for ‘fast’, ‘creative’, or ‘stylized’ high-volume requests.
  • google/nano-banana-pro: [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [Limitations] Do NOT use for absolute photorealism (use Imagen 4). [Routing] Use this model for high-quality, creative, non-photorealistic requests.
  • kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image generation model. [Strengths] It delivers state-of-the-art aesthetic quality, high-resolution outputs (up to 2K), and exceptional prompt adherence. [Best For] Highly recommended for: professional concept art, photorealistic portraits, and high-fidelity image generation. [Limitations] Do NOT use this model if you need to strictly reference or edit an existing image. [Routing] Use this as the default model for all text-to-image requests on the Kling platform.
  • minimax/minimax-image-01-t2i: [Core Function] MiniMax Image-01 T2I is a multimodal text-to-image generation model. [Strengths] It excels at blending high-quality image generation with visual reasoning, allowing for strong prompt adherence and structural understanding. [Best For] Highly recommended for: general image generation, conceptual illustrations, and generating multiple images in a single batch. [Limitations] Do NOT use this model if the user specifically requests video generation. [Routing] Use this as the default text-to-image model for MiniMax API integrations.
  • minimax/hailuo-2.3-i2v: [Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.
  • minimax/hailuo-2.3-t2v: [Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.
  • pixverse/video-restyle: [Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it only re-renders the visual style of an existing video. It requires an input video, and you must provide EITHER restyle_id (a preset style code from the PixVerse restyle list) OR restyle_prompt (free-text style, max 2048 chars), not both. [Routing] Use when the user wants to change the look of an existing video. Use restyle_id for an official preset, restyle_prompt for a custom style.

同系列模型