google/veo-3.1-i2v

veo-3.1-i2v
文档
Schema

[Core Function] Veo 3.1 I2V is Google's cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animating concept art, creating cinematic transitions between images, and high-end video production. [Limitations] When using first+last frame or reference-only modes, the `duration` parameter must strictly be 8. Negative prompts are not supported in reference-only mode. [Routing] Use this model by default for high-quality image-to-video tasks or when multiple reference images are provided.

$0.3360~$0.5000/sec
image-to-video

输入

Video description text guiding the animation
Starting frame. Accepts an image URL (e.g. `https://example.com/frame.jpg`).
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Text describing what to avoid in the generated video.
Video aspect ratio
16:9
Video duration in seconds (string type). Constraint: When using `lastFrame`, must be `8` — values `4` or `6` will be rejected by Google. Otherwise `4`, `6`, `8` are all valid.
duration
End frame. Accepts an image URL. Constraint: `duration` must be `8` when using this parameter.
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Person generation policy
personGeneration
Video resolution. Constraint: 1080p and 4k are only available when `duration` is `8`.
resolution

结果

暂无结果

运行模型后,结果将在这里显示。

Next:

README

参数

参数名 描述 类型 必填 枚举值
prompt Video description text guiding the animation string -
image Starting frame. Accepts an image URL (e.g. https://example.com/frame.jpg). string -
negativePrompt Text describing what to avoid in the generated video. string -
aspectRatio Video aspect ratio string 16:9, 9:16
duration Video duration in seconds (string type). Constraint: When using lastFrame, must be 8 — values 4 or 6 will be rejected by Google. Otherwise 4, 6, 8 are all valid. string 4, 6, 8
lastFrame End frame. Accepts an image URL. Constraint: duration must be 8 when using this parameter. string -
personGeneration Person generation policy string allow_adult
resolution Video resolution. Constraint: 1080p and 4k are only available when duration is 8. string 720p, 1080p, 4k

价格

单位: $/sec

维度 价格
resolution: 720p 0.3360
resolution: 1080p 0.3360
resolution: 4k 0.5000

相关模型

  • google/veo-3.1-fast-t2v: [Core Function] Veo 3.1 Fast T2V is a high-speed text-to-video model. [Strengths] It is heavily optimized for fast generation, delivering video content with natively synchronized audio at 1080p quickly. [Best For] Highly recommended for: rapid prototyping, quick visual iteration, and high-volume background video generation. [Limitations] Do NOT use this model if you require 4K resolution or maximum artistic detail. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or ‘rapid’ video generation.
  • google/veo-3.1-t2v: [Core Function] Veo 3.1 T2V is Google’s state-of-the-art cinematic text-to-video engine. [Strengths] It natively generates 4K professional-grade video output with natively synchronized audio and supports complex camera movements. [Best For] Highly recommended for: high-end creative storytelling, cinematic short films, and experimental video production with sound. [Limitations] Do NOT use this model if you need instant/real-time generation, as 4K video rendering takes time. [Routing] Use this model by default for all high-quality text-to-video requests on the Google platform.
  • google/veo-3.1-fast-i2v: [Core Function] Veo 3.1 Fast I2V is a high-speed image-to-video model. [Strengths] It quickly animates starting images at 1080p, optimized for low latency. [Best For] Highly recommended for: rapid prototyping and quick social media visual iterations. [Limitations] Do NOT use this model for the absolute highest visual fidelity or 4K output. [Routing] Choose this model when the user emphasizes ‘fast’ or ‘quick’ animation.
  • google/veo-3.1-lite-t2v: [Core Function] Veo 3.1 Lite T2V is a balanced text-to-video model. [Strengths] It provides a good balance between generation speed and visual quality, still supporting the advanced architecture of the 3.1 series. [Best For] Highly recommended for: general video content creation and social media posts where 4K is not strictly necessary. [Limitations] Do NOT use this model for the absolute highest cinematic fidelity (use the standard Veo 3.1 instead). [Routing] Route to this model for standard, everyday video generation tasks.
  • google/veo-3.1-lite-i2v: [Core Function] Veo 3.1 Lite I2V is a balanced image-to-video model. [Strengths] It offers a middle ground between speed and quality for animating images. [Best For] Highly recommended for: general image animation and web-ready content. [Limitations] Do NOT use this model if you need 4K resolution. [Routing] Use this model for standard image animation requests.
  • kling/kling-v3-i2v: [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.
  • bytedance/seedance-2.0-i2v: [Core Function] Seedance 2.0 I2V is ByteDance’s flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.
  • minimax/hailuo-2.3-i2v: [Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.
  • vidu/viduq3-pro-i2v: [Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended for: bringing concept art to life, professional film production, high-end commercial showcases, and creating immersive environments from still images. [Limitations] Do NOT use this model if you need instant/real-time generation, as rendering takes longer. It does not support 4K resolution. [Routing] Use this model by default for high-quality image-to-video requests. If the user requires faster generation, route to Q3 Pro Fast or Q3 Turbo.
  • vidu/viduq3-pro-fl2v: [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.
  • skywork/skyreels-i2v: [Core Function] SkyReels Image-to-Video animates one or more keyframe images into a video guided by a text prompt. [Strengths] Supports a start frame, an end frame, and tagged mid-frames for keyframe control; optional audio, 480p/720p/1080p output, and fast/std modes. [Best For] Bringing a photo to life, first-last-frame transitions, and keyframe-driven storyboards. [Limitations] Do NOT use this for pure text-to-video (use skyreels-t2v) or for editing an existing video (use the Omni / video-to-video models). It requires at least one of first_frame_image, end_frame_image, or mid_frame_images; output is capped at 1080p and 15s, and fast mode supports only sound=false. [Routing] Provide first_frame_image to animate from a start image, add end_frame_image for a transition, or supply mid_frame_images (each tag must appear in the prompt as @tag) for keyframe guidance.
  • pixverse/motion-control: [Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limitations] Do NOT use this if you only have a video and no subject image (use Restyle or Extend instead), or if you need 1080p output (only 360p/540p/720p are supported). It requires BOTH a subject image (with a clear person or animal) AND a reference video (with a person as the primary focus). [Routing] Use when the user has one subject image and one motion reference video and wants the subject to mimic that motion.
  • vidu/motion-sync: [Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target’s identity. [Best For] Highly recommended for: creating dance videos with custom characters, transferring complex choreography, and replicating specific physical actions onto avatars. [Limitations] Do NOT use this model if you want to change what a character is saying (use Lip Sync). It requires both a reference video for motion and a target image for appearance. [Routing] Use this model specifically when the user wants to copy the body movements or actions from one video onto a different character.

同系列模型