vidu/viduq2-pro-digital-human

Docs
Schema

[Core Function] Vidu Q2 Pro Digital Human is a premium portrait animation model. [Strengths] It excels at generating highly realistic, expressive digital humans from a single portrait image, featuring precise lip-sync to audio and natural facial micro-expressions. [Best For] Highly recommended for: professional virtual spokespersons, high-end educational videos, news anchoring, and realistic character animation. [Limitations] Do NOT use this model for complex full-body physical interactions or videos longer than 10 seconds. [Routing] Use this by default for 'talking head' or 'digital human' requests prioritizing realism over speed.

$0.1380/sec
image-to-video

Input

Motion or speech description for the digital human
Portrait image URL or base64 data URI
Hint: Drag and drop files, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Audio file URL. When provided, the digital human's lip movements will synchronize to this audio track.
Hint: Drag and drop files, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Video resolution
720p

Result

No results yet

Run the model to preview the output here.

Next:

README

Parameters

Name Description Type Required Enums
prompt Motion or speech description for the digital human string No -
image_url Portrait image URL or base64 data URI string Yes -
audio_url Audio file URL. When provided, the digital human’s lip movements will synchronize to this audio track. string No -
resolution Video resolution string No 540p, 720p, 1080p

Pricing

Unit: $/sec

Pricing
$0.1380/sec
  • vidu/viduq2-turbo-digital-human: [Core Function] Vidu Q2 Turbo Digital Human is a fast portrait animation model. [Strengths] It excels at quickly animating a static portrait image into a speaking or moving digital human, syncing lip movements to provided audio with low latency. [Best For] Highly recommended for: rapid generation of talking head videos, quick virtual presenters, and responsive interactive avatars. [Limitations] Do NOT use this model for complex full-body motion, multi-character interactions, or videos longer than 10 seconds. [Routing] Use this model when the user wants to make a portrait ‘talk’ quickly. For higher realism and better quality, route to Q2 Pro Digital Human.
  • openai/gpt-image-1.5: [Core Function] GPT Image 1.5 is a versatile text-to-image generation model. [Strengths] It balances solid visual performance with crucial utility features, notably its native support for generating images with transparent backgrounds. [Best For] Highly recommended for: creating UI icons, standalone logos, game assets, and any graphic design elements that require a transparent background. [Limitations] Do NOT use this model if you need 2K or 4K resolution. Its maximum supported resolution is 1536x1024. [Routing] Choose this model specifically when the user asks for ‘transparent background’, ‘no background’, or ‘PNG icon’. For standard, high-fidelity, or 4K image generation, use GPT Image 2 instead.
  • google/nano-banana-2: [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism (use the Imagen 4 series instead). [Routing] Route to this model for ‘fast’, ‘creative’, or ‘stylized’ high-volume requests.
  • google/nano-banana-pro: [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [Limitations] Do NOT use for absolute photorealism (use Imagen 4). [Routing] Use this model for high-quality, creative, non-photorealistic requests.
  • kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image generation model. [Strengths] It delivers state-of-the-art aesthetic quality, high-resolution outputs (up to 2K), and exceptional prompt adherence. [Best For] Highly recommended for: professional concept art, photorealistic portraits, and high-fidelity image generation. [Limitations] Do NOT use this model if you need to strictly reference or edit an existing image. [Routing] Use this as the default model for all text-to-image requests on the Kling platform.
  • openai/gpt-image-1.5-edit: [Core Function] GPT Image 1.5 Edit is a versatile image-to-image editing and merging model. [Strengths] It excels at complex utility editing tasks, including multi-image merging (up to 16 images), transparent background support, and precise control over how strictly the model adheres to the input image (fidelity control). [Best For] Highly recommended for: merging reference images, editing UI assets, generating variations with strict shape preservation, and creating transparent cutouts. [Limitations] Do NOT use this model if you require 2K or 4K high-resolution outputs, as it is limited to standard resolutions. [Routing] Use this model specifically when the user provides multiple images to combine, requires transparency, or explicitly asks to ‘keep the exact shape’ of the original image (fidelity control). Otherwise, use GPT Image 2 Edit.
  • kling/kling-avatar: [Core Function] Kling Avatar is a specialized portrait animation model. [Strengths] It precisely animates a portrait image to lip-sync with an audio file or TTS audio ID. [Best For] Highly recommended for: virtual presenters, talking head videos, and digital avatars. [Limitations] Do NOT use this model for full-body action or general image animation. [Routing] Use this model explicitly when the user wants to make a portrait ‘speak’ with provided audio.
  • skywork/single-actor-avatar: [Core Function] SkyReels Single-Actor Avatar (audio-to-video) drives a talking-avatar video from a single portrait image and one audio track. [Strengths] Lip-synced single-speaker talking-head video generated from an image plus audio. [Best For] Virtual presenters, single-speaker dubbing, and talking avatars. [Limitations] Do NOT use this for multi-speaker scenes (use the multi-actor flow) or when you only have text. It requires first_frame_image and exactly one audio segment (<=200s). mode=std outputs 720p, mode=pro outputs 1080p. [Routing] Provide a portrait first_frame_image and one audio URL in audios; choose mode=pro for 1080p output.