Attributes
- Provider: vidu
- Category: Image to Video
Collections
Parameters
| Name | Description | Type | Required | Enums |
|---|---|---|---|---|
| prompt | Marketing message or product description | string | No | - |
| images | Product or scene images (1-7) as URLs or base64 data URIs | string[] | Yes | - |
| aspect_ratio | Video aspect ratio | string | No | 16:9, 9:16, 1:1 |
| duration | Video duration in seconds (10-60) | integer | No | - |
| language | Script language for generated narration | string | No | zh, en |
Pricing
Unit: $/sec
| Pricing |
|---|
| $0.1840/sec |
Related Models
- vidu/one-click-general-film: [Core Function] Vidu One-Click General Film is an automated cinematic film generation model. [Strengths] It excels at automatically stringing together 1 to 7 user-provided images into a cohesive, cinematic film (up to 180s) with appropriate transitions and pacing. [Best For] Highly recommended for: instant music videos, cinematic montages, automated travel vlogs, and turning photo albums into compelling short films. [Limitations] Do NOT use this model if the user needs precise, frame-by-frame control over camera movements or specific character actions in each shot. [Routing] Use this when the user wants an automated ‘done-for-you’ long video from a batch of images without manually prompting every single shot.
- vidu/one-click-trending-replicate: [Core Function] Vidu One-Click Trending Replicate is a viral video style cloning model. [Strengths] It excels at analyzing a trending or viral reference video and recreating its specific visual style, transitions, and pacing using the user’s own provided subject images. [Best For] Highly recommended for: participating in social media video trends, quickly cloning viral visual effects, and applying popular editing styles to personal photos. [Limitations] Do NOT use this model for original cinematic storytelling. It is strictly meant to mimic the style of a provided reference video. [Routing] Use this when the user explicitly provides a ‘trend’ or ‘viral’ video and wants to recreate that exact vibe or transition style with their own images.
- google/gemini-omni-flash-t2v: [Core Function] Gemini Omni Flash T2V is Google’s fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.




