属性
コレクション
パラメーター
| 名前 | 説明 | 型 | 必須 | 列挙値 |
|---|---|---|---|---|
| prompt | Motion or speech description for the digital human | string | いいえ | - |
| image_url | Portrait image URL or base64 data URI | string | はい | - |
| audio_url | Audio file URL. When provided, the digital human’s lip movements will synchronize to this audio track. | string | いいえ | - |
| resolution | Video resolution | string | いいえ | 540p, 720p, 1080p |
料金
単位: $/sec
| 料金 |
|---|
| $0.0050/sec |
関連モデル
- vidu/viduq2-turbo-digital-human: [Core Function] Vidu Q2 Turbo Digital Human is a fast portrait animation model. [Strengths] It excels at quickly animating a static portrait image into a speaking or moving digital human, syncing lip movements to provided audio with low latency. [Best For] Highly recommended for: rapid generation of talking head videos, quick virtual presenters, and responsive interactive avatars. [Limitations] Do NOT use this model for complex full-body motion, multi-character interactions, or videos longer than 10 seconds. [Routing] Use this model when the user wants to make a portrait ‘talk’ quickly. For higher realism and better quality, route to Q2 Pro Digital Human.
- google/nano-banana-2: [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism; prefer Nano Banana Pro for higher-fidelity creative output. [Routing] Route to this model for ‘fast’, ‘creative’, or ‘stylized’ high-volume requests.
- google/nano-banana-pro: [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [Limitations] Do NOT use this model when you need the absolute highest photorealism or detail; step up within the Nano Banana family (Nano Banana 2 or Nano Banana Pro Edit workflows) as needed. [Routing] Use this model for high-quality, creative, non-photorealistic requests.
- kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.
- kling/kling-avatar: [Core Function] Kling Avatar is a specialized portrait animation model. [Strengths] It precisely animates a portrait image to lip-sync with an audio file or TTS audio ID. [Best For] Highly recommended for: virtual presenters, talking head videos, and digital avatars. [Limitations] Do NOT use this model for full-body action or general image animation. [Routing] Use this model explicitly when the user wants to make a portrait ‘speak’ with provided audio.
- skywork/single-actor-avatar: [Core Function] SkyReels Single-Actor Avatar (audio-to-video) drives a talking-avatar video from a single portrait image and one audio track. [Strengths] Lip-synced single-speaker talking-head video generated from an image plus audio. [Best For] Virtual presenters, single-speaker dubbing, and talking avatars. [Limitations] Do NOT use this for multi-speaker scenes (use the multi-actor flow) or when you only have text. It requires first_frame_image and exactly one audio segment (<=200s). mode=std outputs 720p, mode=pro outputs 1080p. [Routing] Provide a portrait first_frame_image and one audio URL in audios; choose mode=pro for 1080p output.
- google/gemini-omni-flash-t2v: [Core Function] Gemini Omni Flash T2V is Google’s fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.








