bytedance/seedance-2.5-i2v

seedance-2.5-i2v
文档
Schema

[Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 480p/720p/1080p, 4-30 second duration, and mp4 or mov output. [Best For] Highly recommended for: animating a keyframe, first-to-last transitions, multi-image character consistency, and image-led storytelling on Seedance 2.5. [Limitations] Do NOT mix first_frame_image or last_frame_image with reference_images. Do NOT send video_urls on I2V. Do NOT use audio_urls alone. Do NOT use this model when the user requires 4k output. [Routing] Prefer this model for Seedance 2.5 image-driven generation. Use Seedance 2.5 T2V for text-only requests and Seedance 2.5 V2V when reference video is required.

$0.1030~$0.5690/sec
image-to-video

输入

Optional. Video description. Maximum 10000 characters. Frame animation description.
Required. Starting frame URL. Do not pass reference_images.
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Optional. Reference audio URLs (1-10). Use with frame images; audio-only is not supported.
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Whether to keep the camera fixed during generation
Video duration in seconds. Allowed values are integers from 4 to 30
Task expiration time in seconds
Whether to generate an audio track with the video
Optional. Ending frame URL. Requires first_frame_image. Do not pass reference_images.
提示:可拖拽文件、从剪贴板粘贴(Ctrl/Cmd+V),或提供 URL。
Output video container. mp4 for general playback; mov for higher color precision in editing workflows
Aspect ratio of the generated video. For first-frame and first-last-frame, use adaptive or omit; a fixed ratio may fail asynchronously with InvalidParameter.TaskTypeConstraint.
Specifies the resolution tier for the generated video. Seedance 2.5 supports 480p, 720p, and 1080p
Whether to return the last frame of the generated video
Random number seed for reproducibility. Use -1 for random. Identical seeds cannot guarantee completely identical results

结果

暂无结果

运行模型后,结果将在这里显示。

Next:

README

属性

系列

集合

参数

参数名 描述 类型 必填 枚举值
prompt Optional. Video description. Maximum 10000 characters. Frame animation description. string -
first_frame_image Required. Starting frame URL. Do not pass reference_images. string -
audio_urls Optional. Reference audio URLs (1-10). Use with frame images; audio-only is not supported. string[] -
camera_fixed Whether to keep the camera fixed during generation boolean true, false
duration Video duration in seconds. Allowed values are integers from 4 to 30 integer 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30
execution_expires_after Task expiration time in seconds integer -
generate_audio Whether to generate an audio track with the video boolean true, false
last_frame_image Optional. Ending frame URL. Requires first_frame_image. Do not pass reference_images. string -
output_format Output video container. mp4 for general playback; mov for higher color precision in editing workflows string mp4, mov
ratio Aspect ratio of the generated video. For first-frame and first-last-frame, use adaptive or omit; a fixed ratio may fail asynchronously with InvalidParameter.TaskTypeConstraint. string 16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
resolution Specifies the resolution tier for the generated video. Seedance 2.5 supports 480p, 720p, and 1080p string 480p, 720p, 1080p
return_last_frame Whether to return the last frame of the generated video boolean true, false
seed Random number seed for reproducibility. Use -1 for random. Identical seeds cannot guarantee completely identical results integer -

价格

单位: $/sec

维度 价格
resolution: 480p 0.1030
resolution: 720p 0.2310
resolution: 1080p 0.5690

相关模型

  • bytedance/seedance-2.0-fast-i2v: [Core Function] Seedance 2.0 Fast I2V is a high-speed multimodal video generation model. [Strengths] Fast generation with the multimodal and multi-shot capabilities of the Seedance 2.0 architecture. [Best For] Highly recommended for: rapid prototyping and quick multi-shot video creation. [Limitations] Do NOT use this model for the absolute highest cinematic fidelity (use the standard 2.0 model instead). [Routing] Choose this model when the user emphasizes ‘fast’ or ‘quick’ generation.
  • bytedance/seedance-2.0-i2v: [Core Function] Seedance 2.0 I2V is ByteDance’s flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.
  • bytedance/seedance-2.0-fast-v2v: [Core Function] Seedance 2.0 Fast V2V is a high-speed video-to-video generation model. [Strengths] It offers rapid video transformation capabilities based on the 2.0 architecture. [Best For] Highly recommended for: quick video restyling and fast iterations. [Limitations] Do NOT use this model when absolute maximum visual quality is required (use standard 2.0). [Routing] Use this when speed is the primary concern for video transformations.
  • bytedance/seedance-2.0-v2v: [Core Function] Seedance 2.0 V2V is ByteDance’s flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] Highly recommended for: complex video-to-video transformations, restyling existing footage, and creating dynamic multi-shot edits. [Limitations] Do NOT use this model for simple still-image generation (use Seedream instead). [Routing] Use this model by default for any video editing or video-to-video generation tasks.
  • bytedance/seedance-2.0-fast-t2v: [Core Function] Seedance 2.0 Fast T2V is the faster text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use when speed is preferred over maximum resolution.
  • bytedance/seedance-2.0-mini-t2v: [Core Function] Seedance 2.0 Mini T2V is the lightweight text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use for cost-efficient text-to-video.
  • bytedance/seedance-2.0-t2v: [Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.
  • bytedance/seedance-2.0-mini-i2v: [Core Function] Seedance 2.0 Mini I2V is the lightweight multimodal image-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with optional first/last frame, reference images, and audio references. [Routing] Use for cost-efficient image-to-video.
  • bytedance/seedance-2.0-mini-v2v: [Core Function] Seedance 2.0 Mini V2V is the lightweight video-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with required reference video and optional text/image/audio references. [Routing] Use for cost-efficient video-to-video transformations.
  • bytedance/seedance-2.5-t2v: [Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p/1080p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-to-video storytelling, social clips beyond 15 seconds, and Seedance workflows that need mov output. [Limitations] Do NOT use this model if the user provides images, video, or audio as inputs. Do NOT use this model when the user requires 4k output. [Routing] Prefer Seedance 2.5 T2V when the user needs more than 15 seconds of text-to-video. Use Seedance 2.0 T2V when 4k is required. Use Seedance 2.5 I2V or V2V when media inputs are provided.
  • bytedance/seedance-2.5-v2v: [Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 480p/720p/1080p, 4-30 second output, and mp4 or mov containers. [Best For] Highly recommended for: editing existing clips, extending motion from a base video, multi-reference restyling, and prompt-driven composition that cites @Video1 or @Image1. [Limitations] Do NOT use this model for text-only or first-frame-only workflows. video_urls is required. Do NOT send first_frame_image or last_frame_image. Do NOT use this model when the user requires 4k output. Do NOT use this model for audio-only input. [Routing] Prefer this model for Seedance 2.5 edit, extension, and video-reference jobs. Use Seedance 2.5 T2V for text-only and Seedance 2.5 I2V for image-first generation.
  • kling/kling-v3-i2v: [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.
  • google/veo-3.1-i2v: [Core Function] Veo 3.1 I2V is Google’s cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animating concept art, creating cinematic transitions between images, and high-end video production. [Limitations] When using first+last frame or reference-only modes, the duration parameter must strictly be 8. Negative prompts are not supported in reference-only mode. [Routing] Use this model by default for high-quality image-to-video tasks or when multiple reference images are provided.
  • minimax/hailuo-2.3-i2v: [Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.
  • vidu/viduq3-pro-i2v: [Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended for: bringing concept art to life, professional film production, high-end commercial showcases, and creating immersive environments from still images. [Limitations] Do NOT use this model if you need instant/real-time generation, as rendering takes longer. It does not support 4K resolution. [Routing] Use this model by default for high-quality image-to-video requests. If the user requires faster generation, route to Q3 Pro Fast or Q3 Turbo.
  • vidu/viduq3-pro-fl2v: [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.
  • skywork/skyreels-i2v: [Core Function] SkyReels Image-to-Video animates one or more keyframe images into a video guided by a text prompt. [Strengths] Supports a start frame, an end frame, and tagged mid-frames for keyframe control; optional audio, 480p/720p/1080p output, and fast/std modes. [Best For] Bringing a photo to life, first-last-frame transitions, and keyframe-driven storyboards. [Limitations] Do NOT use this for pure text-to-video (use skyreels-t2v) or for editing an existing video (use the Omni / video-to-video models). It requires at least one of first_frame_image, end_frame_image, or mid_frame_images; output is capped at 1080p and 15s, and fast mode supports only sound=false. [Routing] Provide first_frame_image to animate from a start image, add end_frame_image for a transition, or supply mid_frame_images (each tag must appear in the prompt as @tag) for keyframe guidance.
  • pixverse/motion-control: [Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limitations] Do NOT use this if you only have a video and no subject image (use Restyle or Extend instead), or if you need 1080p output (only 360p/540p/720p are supported). It requires BOTH a subject image (with a clear person or animal) AND a reference video (with a person as the primary focus). [Routing] Use when the user has one subject image and one motion reference video and wants the subject to mimic that motion.
  • vidu/motion-sync: [Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target’s identity. [Best For] Highly recommended for: creating dance videos with custom characters, transferring complex choreography, and replicating specific physical actions onto avatars. [Limitations] Do NOT use this model if you want to change what a character is saying (use Lip Sync). It requires both a reference video for motion and a target image for appearance. [Routing] Use this model specifically when the user wants to copy the body movements or actions from one video onto a different character.
  • google/gemini-omni-flash-t2v: [Core Function] Gemini Omni Flash T2V is Google’s fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.
  • google/gemini-omni-flash-i2v: [Core Function] Gemini Omni Flash I2V is a fast Image-to-Video model that animates a single input image into a short 720p video via the Interactions API. [Strengths] It uses the provided image as the opening frame and generates smooth motion with natively synchronized audio at low latency. [Best For] Highly recommended for: bringing a still photo to life, quick product or portrait animation, and short social clips derived from a single image. [Limitations] Do NOT use this model if you need 1080p or 4K output, clips longer than 10 seconds, or the fusion of multiple reference images; it takes exactly one image and outputs 720p up to 10 seconds (16:9 or 9:16). Do NOT use it to edit an existing video. [Routing] Choose this when the user provides one image to animate. To fuse multiple reference images use Gemini Omni Flash R2V; to edit an existing video use Gemini Omni Flash Video Edit; for 4K cinematic results use Veo 3.1 I2V.
  • google/gemini-omni-flash-r2v: [Core Function] Gemini Omni Flash R2V (Reference-to-Video) generates a short 720p video guided by up to three reference images via the Interactions API. [Strengths] It fuses the styles, subjects, or elements from multiple reference images (referred to in the text prompt) into a single coherent animated clip with synchronized audio. [Best For] Highly recommended for: blending characters or visual styles from several images, reference-guided creative shots, and multi-subject compositions where the prompt directs how the references combine. [Limitations] Do NOT use this model if you only have a single starting frame (use I2V instead), or if you need 1080p or 4K or clips longer than 10 seconds; it accepts 1 to 3 reference images and outputs 720p up to 10 seconds (16:9 or 9:16). [Routing] Choose this when the user supplies multiple reference images to combine into one video. For single first-frame animation use Gemini Omni Flash I2V; to modify an existing video use Gemini Omni Flash Video Edit.
  • kling/kling-v3-omni-i2v: [Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.
  • kling/kling-v3-turbo-i2v: [Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.
  • alibaba/wan3.0-i2v: [Core Function] Wan 3.0 I2V is Alibaba Wan 3.0 image-to-video generation supporting first-frame, first-last-frame, and reference-image modes. [Strengths] It can strictly lock the first and last frames or fuse up to 10 reference images with optional reference audio for multimodal guidance. [Best For] Highly recommended for: animating a single keyframe, cinematic first-to-last transitions, multi-image character or product consistency, and image-led storytelling. [Limitations] Do NOT use this model for text-only generation, document/webpage reference, or when a reference video is required. Do NOT mix first_frame/last_frame with reference_images/audio_urls. [Routing] Prefer this model for Wan 3.0 image-driven video. Use Wan 3.0 T2V for prompt/file/link inputs and Wan 3.0 V2V when video_urls are provided.

相关资源