属性
シリーズ
コレクション
パラメーター
| 名前 | 説明 | 型 | 必須 | 列挙値 |
|---|---|---|---|---|
| prompt | Text description of the edit to apply to the video | string | はい | - |
| video | Input video URL or base64 data URI | string | はい | - |
料金
単位: $/sec
| 料金 |
|---|
| $0.1725/sec |
関連モデル
- xai/grok-imagine-image: [Core Function] Grok Imagine Image is xAI’s standard text-to-image generation model. [Strengths] It excels at quickly generating solid, visually appealing images from a text prompt across a wide range of aspect ratios. [Best For] Highly recommended for: rapid prototyping, social media content, and general-purpose image generation. [Limitations] Do NOT use this model when maximum detail or fidelity is required; the Quality variant produces richer detail. [Routing] Choose this model for fast, general image generation. When the user demands maximum fidelity, route to Grok Imagine Image (Quality).
- xai/grok-imagine-image-quality: [Core Function] Grok Imagine Image (Quality) is xAI’s high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed concept art, marketing visuals, and any scenario where image quality is prioritized over generation speed. [Limitations] Do NOT use this model when latency is critical, as generation is slower than the standard model. [Routing] Use this model by default when the user emphasizes quality or detail. For faster, lighter generation use Grok Imagine Image (standard).
- xai/grok-imagine-video: [Core Function] Grok Imagine Video is xAI’s text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social clips, animated concepts, and dynamic scene generation from a description. [Limitations] Do NOT use this model when you have a starting image or reference subjects, or when you need resolutions above 720p or clips longer than 15 seconds; it is limited to 480p/720p and 15s. [Routing] Use this model when the user wants a video from text only. If a starting image is provided, route to the Image-to-Video model; for reference-driven character video, use Reference-to-Video.
- xai/grok-imagine-image-edit: [Core Function] Grok Imagine Image Edit is xAI’s standard image editing model. [Strengths] It excels at quickly applying prompt-guided edits and style changes to one or more source images. [Best For] Highly recommended for: fast restyling, quick variations, and lightweight image edits. [Limitations] Do NOT use this model when maximum edit fidelity is required; the Quality variant preserves more detail. A maximum of 3 source images is supported. [Routing] Choose this model for fast edits. When the user demands maximum fidelity, route to Grok Imagine Image Edit (Quality).
- xai/grok-imagine-image-quality-edit: [Core Function] Grok Imagine Image Edit (Quality) is xAI’s high-fidelity image editing model. [Strengths] It excels at applying detailed, prompt-guided edits and style transformations to one or more source images while preserving fine detail. [Best For] Highly recommended for: high-quality restyling, detailed inpainting-style edits, and combining up to 3 source images. [Limitations] Do NOT use this model when latency is critical, as it is slower than the standard edit variant; a maximum of 3 source images is supported. [Routing] Use this model by default for quality-sensitive edits. For faster edits, route to Grok Imagine Image Edit (standard).
- xai/grok-imagine-video-1.5-i2v: [Core Function] Grok Imagine Video 1.5 I2V animates a single starting image into a video using the Grok Imagine 1.5 generation backbone. [Strengths] It excels at producing motion from one starting frame with the improved 1.5 model. [Best For] Highly recommended for: animating a photo or illustration when the 1.5 generation backbone is preferred. [Limitations] Do NOT use this model for text-only generation, for resolutions above 1080p, or for clips longer than 15 seconds; it requires a starting image and is limited to 480p/720p/1080p and 15s. [Routing] Use this model when the user provides one starting image and prefers the 1.5 backbone.
- xai/grok-imagine-video-i2v: [Core Function] Grok Imagine Video I2V animates a single starting image into a video. [Strengths] It excels at producing smooth motion from one starting frame, guided by a text prompt for the desired movement. [Best For] Highly recommended for: bringing a photo or illustration to life, dynamic product showcases, and quick cinematic motion from a still. [Limitations] Do NOT use this model for text-only generation, for resolutions above 720p, or for clips longer than 15 seconds; it requires a starting image and is limited to 480p/720p and 15s. [Routing] Use this model when the user provides exactly one starting image. For reference-driven character video, use Reference-to-Video; for text-only video, use Text-to-Video.
- xai/grok-imagine-video-r2v: [Core Function] Grok Imagine Video R2V generates a video from a text prompt while preserving the subjects shown in up to 7 reference images. [Strengths] It excels at keeping character/subject identity consistent across a newly generated scene driven by the prompt. [Best For] Highly recommended for: character-driven video, placing a specific subject into a new scene, and blending features from multiple references. [Limitations] Do NOT use this model to simply animate a single image as-is (use Image-to-Video), or for clips longer than 10 seconds or resolutions above 720p; duration is capped at 10s and resolution at 480p/720p. [Routing] Use this model when the user provides reference images of a subject and wants a new action/scene described by a prompt. To simply animate a single image as-is, use Image-to-Video.
- xai/grok-imagine-video-extend: [Core Function] Grok Imagine Video Extend continues an existing video, generating additional footage beyond its end. [Strengths] It excels at seamlessly extending a clip with new prompt-guided motion. [Best For] Highly recommended for: lengthening short clips, continuing a scene, and adding follow-on action. [Limitations] Do NOT use the
durationparameter expecting it to set the total video length; it only controls the length of the appended segment (2-10 seconds). Input video constraints are enforced by the upstream provider. [Routing] Use this model when the user wants to make a video longer. To restyle or modify an existing video, use Video Edit. - alibaba/happyhorse-1.1-t2v: [Core Function] HappyHorse 1.1 T2V is Alibaba’s latest streamlined text-to-video model. [Strengths] It generates 720P/1080P video with native audio support, 3-15 second duration, and an expanded set of aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: fast HappyHorse text-to-video generation, social video formats, and high-throughput content creation. [Limitations] It does not expose custom audio controls; use Wan 2.7 T2V when custom audio input is required. [Routing] Prefer this model when the user explicitly requests HappyHorse text-to-video or wants the latest HappyHorse generation quality.
- alibaba/wan2.7-t2v: [Core Function] Wan 2.7 T2V is Alibaba’s flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: high-quality commercial video generation, professional storytelling, and dynamic cinematic sequences. [Limitations] Do NOT use this model if the user specifically requests the streamlined ‘HappyHorse’ workflow. [Routing] Use this model by default for high-end text-to-video requests on the Alibaba platform.
- bytedance/seedance-2.0-t2v: [Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.
- bytedance/seedance-2.0-i2v: [Core Function] Seedance 2.0 I2V is ByteDance’s flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.
- alibaba/happyhorse-1.1-i2v: [Core Function] HappyHorse 1.1 I2V is Alibaba’s latest streamlined first-frame image-to-video model. [Strengths] It turns a single image into high-quality 720P/1080P video with native audio support and 3-15 second duration; output aspect ratio follows the first frame image. [Best For] Highly recommended for: rapid image animation, product motion previews, and simple character or scene animation. [Limitations] It does not accept an explicit ratio parameter; use T2V or R2V when you need a fixed generated aspect ratio. [Routing] Prefer this model when the user provides one image and requests HappyHorse image animation.
- alibaba/happyhorse-1.1-r2v: [Core Function] HappyHorse 1.1 R2V is Alibaba’s latest reference-image-to-video model. [Strengths] It uses 1-9 reference images to preserve subject or character appearance while generating new video actions, supports 720P/1080P output, 3-15 second duration, and expanded aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: character-consistent storytelling, reference-based product shots, and multi-image subject composition. [Limitations] Do NOT use if the user simply wants to animate a single image exactly as provided; use HappyHorse I2V instead. [Routing] Use when the user provides one or more reference images and asks for a newly generated HappyHorse video.
- alibaba/wan2.7-i2v: [Core Function] Wan 2.7 I2V is Alibaba’s flagship multimodal image-to-video model. [Strengths] It supports multimodal input (text, image, audio, video) for first-frame, start-and-end-frame (FL2V), and video continuation tasks. [Best For] Highly recommended for: complex image animation, cinematic transitions, and video extension workflows. [Limitations] Do NOT use this model if you only need a quick, simple animation where HappyHorse might be faster. [Routing] Use this model by default for complex image-to-video or video continuation tasks.
- alibaba/wan2.7-r2v: [Core Function] Wan 2.7 Reference-to-Video is a highly capable character/entity reference video model. [Strengths] It natively supports entity reference, voice customization, and playbook-based video generation from a single storyboard. [Best For] Highly recommended for: creating consistent video series, brand mascot animation, and storyboard-driven storytelling. [Limitations] Do NOT use this model for simple, single-image direct animation (use I2V instead). [Routing] Use this by default for complex character consistency and storyboard generation tasks on Alibaba.
- alibaba/wan2.7-videoedit: [Core Function] Wan 2.7 Video Editing is an instruction-based video modification model. [Strengths] It supports complex video editing tasks like content replacement using reference images, and replicating actions, effects, and camera movements. [Best For] Highly recommended for: modifying existing video footage, style transfer on videos, and targeted element replacement. [Limitations] Do NOT use this model to generate a brand new video from scratch; it requires an input video. [Routing] Use this model by default whenever a user wants to edit, alter, or restyle an existing video.
- bytedance/seedance-2.0-v2v: [Core Function] Seedance 2.0 V2V is ByteDance’s flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] Highly recommended for: complex video-to-video transformations, restyling existing footage, and creating dynamic multi-shot edits. [Limitations] Do NOT use this model for simple still-image generation (use Seedream instead). [Routing] Use this model by default for any video editing or video-to-video generation tasks.
- google/gemini-omni-flash-t2v: [Core Function] Gemini Omni Flash T2V is Google’s fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.
- google/gemini-omni-flash-i2v: [Core Function] Gemini Omni Flash I2V is a fast Image-to-Video model that animates a single input image into a short 720p video via the Interactions API. [Strengths] It uses the provided image as the opening frame and generates smooth motion with natively synchronized audio at low latency. [Best For] Highly recommended for: bringing a still photo to life, quick product or portrait animation, and short social clips derived from a single image. [Limitations] Do NOT use this model if you need 1080p or 4K output, clips longer than 10 seconds, or the fusion of multiple reference images; it takes exactly one image and outputs 720p up to 10 seconds (16:9 or 9:16). Do NOT use it to edit an existing video. [Routing] Choose this when the user provides one image to animate. To fuse multiple reference images use Gemini Omni Flash R2V; to edit an existing video use Gemini Omni Flash Video Edit; for 4K cinematic results use Veo 3.1 I2V.
- google/gemini-omni-flash-r2v: [Core Function] Gemini Omni Flash R2V (Reference-to-Video) generates a short 720p video guided by up to three reference images via the Interactions API. [Strengths] It fuses the styles, subjects, or elements from multiple reference images (referred to in the text prompt) into a single coherent animated clip with synchronized audio. [Best For] Highly recommended for: blending characters or visual styles from several images, reference-guided creative shots, and multi-subject compositions where the prompt directs how the references combine. [Limitations] Do NOT use this model if you only have a single starting frame (use I2V instead), or if you need 1080p or 4K or clips longer than 10 seconds; it accepts 1 to 3 reference images and outputs 720p up to 10 seconds (16:9 or 9:16). [Routing] Choose this when the user supplies multiple reference images to combine into one video. For single first-frame animation use Gemini Omni Flash I2V; to modify an existing video use Gemini Omni Flash Video Edit.
- google/gemini-omni-flash-video-edit: [Core Function] Gemini Omni Flash Video Edit performs conversational, instruction-driven editing of an existing video via the Interactions API. [Strengths] It applies natural-language edits (changing the scene, mood, style, lighting, background, or time of day) to an input video while preserving the source video’s length and aspect ratio, with synchronized audio. [Best For] Highly recommended for: re-styling or re-lighting an existing clip, changing a video’s setting or atmosphere, and quick instruction-based revisions of a short video. [Limitations] Do NOT use this model to generate a video from scratch (use T2V, I2V, or R2V), and do NOT expect to change the output resolution, aspect ratio, or duration: the output preserves the source video’s aspect ratio and length, and the model does not accept aspectRatio or duration parameters. The source video should be 3 to 10 seconds. [Routing] Choose this only when the user provides an existing video to modify. To create a new video from text or images, use Gemini Omni Flash T2V, I2V, or R2V instead.
- kling/kling-v3-omni-t2v: [Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and better adherence when scenes demand coherent characters or multi-beat storytelling from text alone. [Best For] Highly recommended for: narrative T2V with recurring subjects, dialogue-aware scenes, brand or product continuity across beats, and premium short films where consistency matters more than raw throughput. [Limitations] Do NOT use this model if the user only needs the cheapest or fastest clip; prefer Kling V3 Turbo T2V. Do NOT use it when the workflow is image-first or needs multi-image references; use Kling V3 Omni I2V or Kling V3 I2V instead. Do NOT use it for deep physics-reasoning specialty tasks better served by Kling Video O1. [Routing] Choose Kling V3 Omni T2V when the user emphasizes Omni, consistency, multimodal quality, or complex text narratives. Prefer Kling V3 T2V as the default high-quality T2V baseline; prefer Kling V3 Turbo T2V when the user stresses speed, cost, or high-volume short-form output.
- kling/kling-v3-t2v: [Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.
- kling/kling-v3-turbo-t2v: [Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical 720p/1080p short videos rather than maximum cinematic headroom. [Best For] Highly recommended for: rapid prototyping, social and ad iteration, batch short-form pipelines, and dialogue clips where turnaround time and unit cost matter most. [Limitations] Do NOT use this model if the user requires peak 4K cinematic fidelity, heavy multi-shot storyboard control, or maximum visual polish; use Kling V3 T2V or Kling V3 Omni T2V instead. Do NOT use it for image-conditioned animation; use Kling V3 Turbo I2V or Kling V3 I2V. [Routing] Choose Kling V3 Turbo T2V when the user says fast, quick, cheap, or high volume. Otherwise default to Kling V3 T2V for quality, or Kling V3 Omni T2V when consistency and Omni-class control are requested.
- kling/kling-v3-i2v: [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.
- kling/kling-v3-omni-i2v: [Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.
- kling/kling-v3-turbo-i2v: [Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.
- kling/kling-v3-omni-video: [Core Function] Kling V3 Omni Video V2V is a multimodal video-to-video endpoint that edits or restyles existing footage using prompt plus optional image and video references. [Strengths] It focuses on source fidelity and subject consistency for Omni-style edit workflows, combining prompt guidance with images and videos inputs so changes stay grounded in the original clip. [Best For] Highly recommended for: reference-faithful video edits, restyling existing takes, keeping characters or products consistent while changing motion or scene instructions, and short-form post workflows that start from real footage. [Limitations] Do NOT use this model for pure text-to-video from scratch; use Kling V3 T2V, Kling V3 Omni T2V, or Kling V3 Turbo T2V. Do NOT use it when you only have a still image and no source video; use Kling V3 I2V or Kling V3 Omni I2V. Do NOT use it for deep physics-reasoning generation better served by Kling Video O1. [Routing] Choose Kling V3 Omni Video V2V when the user already has video to edit or transform and mentions Omni or multimodal references. Prefer Kling Video O1 when reasoning-heavy generation is the goal rather than source-based editing.
- bytedance/seedream-5.0-lite-edit: [Core Function] Seedream 5.0 Lite Edit is a reasoning-enhanced, smart image editing model. [Strengths] It features superior cross-modal understanding and reasoning, allowing for highly accurate, interactive multi-turn image editing with real-time knowledge enhancement. [Best For] Highly recommended for: complex image editing tasks, structural modifications, and edits requiring deep semantic understanding. [Limitations] As a ‘Lite’ model, raw visual rendering might not match the 4.5 tier. [Routing] Use this model by default for complex, reasoning-based image editing tasks.
- kling/kling-v3-i2i: [Core Function] Kling V3 I2I is the flagship image-to-image editing model (POST /images/generations, model_name=kling-v3 with image). [Strengths] High-quality style transfer and editing up to 2K. [Best For] Single-reference image editing. [Limitations] Do NOT send negative_prompt when image is present (officially unsupported). No image_fidelity / image_reference on V3. For multi-image fusion or series, use kling-v3-omni-image. [Routing] Default for standard image-to-image.
- alibaba/wan2.7-image-pro-edit: [Core Function] Wan 2.7 Image Pro Edit is Alibaba’s flagship reasoning-enhanced image editing model. [Strengths] It supports interactive editing, character-consistent multi-image generation, and complex multi-reference modifications with deep reasoning. [Best For] Highly recommended for: professional image retouching, consistent character sheets, and complex structural edits. [Limitations] Do NOT use this model if you specifically need to use negative prompts to exclude elements during editing (use Qwen Image 2.0 Pro Edit instead). [Routing] Use this by default for high-end image editing and multi-reference consistent character generation.
- minimax/minimax-image-01-i2i: [Core Function] MiniMax Image-01 I2I is an image-to-image editing and variation model. [Strengths] It excels at generating new images based on a text prompt while structurally referencing one or more input images. [Best For] Highly recommended for: style transfer, generating variations of existing artwork, and structurally guided image creation. [Limitations] Do NOT use this model if you want to generate video or if you do not have a reference image. [Routing] Use this model when the user provides a reference image and a text prompt to generate a new image.
- google/nano-banana-2-edit: [Core Function] Nano Banana 2 Edit is a high-speed image editing model. [Strengths] It rapidly modifies existing images or extracts image frames from videos based on text prompts. [Best For] Highly recommended for: rapid style transfer, quick image modifications, and fast creative edits. [Limitations] Do NOT use this model for meticulous photorealistic retouching. [Routing] Use this model by default for fast, creative image editing tasks.
- pixverse/video-restyle: [Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it only re-renders the visual style of an existing video. It requires an input video, and you must provide EITHER restyle_id (a preset style code from the PixVerse restyle list) OR restyle_prompt (free-text style, max 2048 chars), not both. [Routing] Use when the user wants to change the look of an existing video. Use restyle_id for an official preset, restyle_prompt for a custom style.
- bytedance/seedream-5.0-pro-edit: [Core Function] Seedream 5.0 Pro Edit is a professional-grade single-image editing (I2I) model. [Strengths] It supports interactive precise editing: edit locations can be specified via coordinates, selection boxes, or arrows described in the prompt, with strong element-level control and subject consistency. [Best For] Highly recommended for: precise local retouching, adding/removing/replacing objects at exact positions, style transfer of a single photo, and professional post-editing workflows. [Limitations] Do NOT use this model for text-to-image generation (an input image is required) or for blending multiple reference images; it accepts exactly one input image and outputs exactly one image (no batch or streaming). [Routing] For 2-10 reference images use Seedream 5.0 Pro Multi-Reference; for pure text-to-image use Seedream 5.0 Pro; choose Seedream 5.0 Lite Edit when batch outputs or more than 10 input images are required.
- bytedance/seedream-5.0-pro-multi-reference: [Core Function] Seedream 5.0 Pro Multi-Reference is a professional-grade multi-reference image generation (I2I) model that creates a single image from 2-10 reference images plus a text prompt. [Strengths] It excels at reference consistency, preserving characters, styles, and objects across multiple input images while following complex blending instructions with professional-grade quality. [Best For] Highly recommended for: keeping character or style consistency across references, combining subjects from different images into one scene, placing products into reference scenes, and IP-consistent content creation. [Limitations] Do NOT use this model with fewer than 2 or more than 10 reference images, and do NOT use it for batch generation or streaming; it outputs exactly one image per request. [Routing] For single-image editing use Seedream 5.0 Pro Edit; for text-only generation use Seedream 5.0 Pro; choose Seedream 5.0 Lite Edit when up to 14 reference images or batch outputs are needed.
- alibaba/qwen-image-3.0-pro-edit: [Core Function] Qwen Image 3.0 Pro Edit is an image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (widthheight), intelligent prompt rewrite, 1-6 outputs while preserving subject identity, and long structured edit instructions. [Best For] Highly recommended for: background replacement, outfit or style changes, multi-image fusion, identity-preserving portrait edits, and iterative creative retouching. [Limitations] Do NOT use this for pure text-to-image with no reference images (use Qwen Image 3.0 Pro instead), native 4K output, or thinking-mode reasoning. Keep output total pixels within 512512 to 2048*2048; input images should follow supported formats and size guidance. Do NOT combine very long prompts with multiple reference images and a long negative_prompt if the request may exceed the model input capacity (about 4.5k tokens total across text and images). [Routing] Route here when the user provides reference image(s) and wants Qwen 3.0 edit quality. For text-only generation without images, use Qwen Image 3.0 Pro.








