vidu/viduq2-turbo-digital-human

Docs
Schema

[Core Function] Vidu Q2 Turbo Digital Human is a fast portrait animation model. [Strengths] It excels at quickly animating a static portrait image into a speaking or moving digital human, syncing lip movements to provided audio with low latency. [Best For] Highly recommended for: rapid generation of talking head videos, quick virtual presenters, and responsive interactive avatars. [Limitations] Do NOT use this model for complex full-body motion, multi-character interactions, or videos longer than 10 seconds. [Routing] Use this model when the user wants to make a portrait 'talk' quickly. For higher realism and better quality, route to Q2 Pro Digital Human.

$0.0920/sec
image-to-video

Input

Motion or speech description for the digital human
Portrait image URL or base64 data URI
Hint: Drag and drop files, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Audio file URL. When provided, the digital human's lip movements will synchronize to this audio track.
Hint: Drag and drop files, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Video resolution
720p

Result

No results yet

Run the model to preview the output here.

Next:

README

Attributes

Collections

Parameters

Name Description Type Required Enums
prompt Motion or speech description for the digital human string No -
image_url Portrait image URL or base64 data URI string Yes -
audio_url Audio file URL. When provided, the digital human’s lip movements will synchronize to this audio track. string No -
resolution Video resolution string No 540p, 720p, 1080p

Pricing

Unit: $/sec

Pricing
$0.0920/sec
  • vidu/viduq2-pro-digital-human: [Core Function] Vidu Q2 Pro Digital Human is a premium portrait animation model. [Strengths] It excels at generating highly realistic, expressive digital humans from a single portrait image, featuring precise lip-sync to audio and natural facial micro-expressions. [Best For] Highly recommended for: professional virtual spokespersons, high-end educational videos, news anchoring, and realistic character animation. [Limitations] Do NOT use this model for complex full-body physical interactions or videos longer than 10 seconds. [Routing] Use this by default for ‘talking head’ or ‘digital human’ requests prioritizing realism over speed.

Related Resources