属性
コレクション
パラメーター
| 名前 | 説明 | 型 | 必須 | 列挙値 |
|---|---|---|---|---|
| prompt | Additional description for the replicated video | string | いいえ | - |
| images | Subject images (1-7) as URLs or base64 data URIs | string[] | はい | - |
| video_url | Reference trending video URL | string | はい | - |
| aspect_ratio | Video aspect ratio | string | いいえ | 16:9, 9:16, 1:1, 4:3, 3:4 |
| remove_audio | Remove audio from the output video | boolean | いいえ | true, false |
| resolution | Video resolution | string | いいえ | 540p, 720p, 1080p |
料金
単位: $/sec
| 条件 | 料金 |
|---|---|
| resolution: 540p | 0.0552 |
| resolution: 720p | 0.0736 |
| resolution: 1080p | 0.0920 |
関連モデル
- vidu/one-click-ad-film: [Core Function] Vidu One-Click AD-Film is an automated marketing video generation model. [Strengths] It excels at transforming 1 to 7 product or scene images into a polished, commercial-style advertisement video (10-60s) automatically. [Best For] Highly recommended for: e-commerce product showcases, social media ads, promotional reels, and quick marketing campaigns. [Limitations] Do NOT use this model for narrative storytelling or cinematic films; it is optimized specifically for commercial pacing and product emphasis. [Routing] Use this specifically when the user wants to generate an ‘ad’, ‘commercial’, or ‘promotional video’ from product photos.
- vidu/one-click-general-film: [Core Function] Vidu One-Click General Film is an automated cinematic film generation model. [Strengths] It excels at automatically stringing together 1 to 7 user-provided images into a cohesive, cinematic film (up to 180s) with appropriate transitions and pacing. [Best For] Highly recommended for: instant music videos, cinematic montages, automated travel vlogs, and turning photo albums into compelling short films. [Limitations] Do NOT use this model if the user needs precise, frame-by-frame control over camera movements or specific character actions in each shot. [Routing] Use this when the user wants an automated ‘done-for-you’ long video from a batch of images without manually prompting every single shot.
- google/gemini-omni-flash-t2v: [Core Function] Gemini Omni Flash T2V is Google’s fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.




