属性
シリーズ
コレクション
パラメーター
| 名前 | 説明 | 型 | 必須 | 列挙値 |
|---|---|---|---|---|
| prompt | Image generation prompt, supports Chinese and English. May reference images as <<<image_1>>> style placeholders. | string | はい | - |
| series_amount | Number of images when result_type is series. Accepts 2–9 or auto. Invalid when result_type is single. |
string | いいえ | - |
| aspect_ratio | Image aspect ratio. When omitted, the platform fills 1:1. |
string | いいえ | 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3, 21:9 |
| images | Optional reference image list (URL or Base64). When provided, must contain 1–10 items. Empty array is invalid. Element library IDs are not exposed. | string[] | いいえ | - |
| n | Number of images to generate when result_type is single. Invalid when result_type is series. |
integer | いいえ | - |
| resolution | Output resolution. Kling V3 Omni supports 1K, 2K, and 4K output. | string | いいえ | 1k, 2k, 4k |
| result_type | Generate a single image or a series. When series, n is invalid; when single, series_amount is invalid. |
string | いいえ | single, series |
料金
単位: $/img
| 条件 | 料金 |
|---|---|
| resolution: 1k | 0.1932 |
| resolution: 2k | 0.1932 |
| resolution: 4k | 0.0386 |
関連モデル
- kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.
- kling/kling-v3-t2v: [Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.
- kling/kling-v3-i2i: [Core Function] Kling V3 I2I is the flagship image-to-image editing model (POST /images/generations, model_name=kling-v3 with image). [Strengths] High-quality style transfer and editing up to 2K. [Best For] Single-reference image editing. [Limitations] Do NOT send negative_prompt when image is present (officially unsupported). No image_fidelity / image_reference on V3. For multi-image fusion or series, use kling-v3-omni-image. [Routing] Default for standard image-to-image.
- kling/kling-v3-i2v: [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.
- kling/kling-v3-omni-video: [Core Function] Kling V3 Omni Video V2V is a multimodal video-to-video endpoint that edits or restyles existing footage using prompt plus optional image and video references. [Strengths] It focuses on source fidelity and subject consistency for Omni-style edit workflows, combining prompt guidance with images and videos inputs so changes stay grounded in the original clip. [Best For] Highly recommended for: reference-faithful video edits, restyling existing takes, keeping characters or products consistent while changing motion or scene instructions, and short-form post workflows that start from real footage. [Limitations] Do NOT use this model for pure text-to-video from scratch; use Kling V3 T2V, Kling V3 Omni T2V, or Kling V3 Turbo T2V. Do NOT use it when you only have a still image and no source video; use Kling V3 I2V or Kling V3 Omni I2V. Do NOT use it for deep physics-reasoning generation better served by Kling Video O1. [Routing] Choose Kling V3 Omni Video V2V when the user already has video to edit or transform and mentions Omni or multimodal references. Prefer Kling Video O1 when reasoning-heavy generation is the goal rather than source-based editing.
- kling/kling-v3-omni-t2v: [Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and better adherence when scenes demand coherent characters or multi-beat storytelling from text alone. [Best For] Highly recommended for: narrative T2V with recurring subjects, dialogue-aware scenes, brand or product continuity across beats, and premium short films where consistency matters more than raw throughput. [Limitations] Do NOT use this model if the user only needs the cheapest or fastest clip; prefer Kling V3 Turbo T2V. Do NOT use it when the workflow is image-first or needs multi-image references; use Kling V3 Omni I2V or Kling V3 I2V instead. Do NOT use it for deep physics-reasoning specialty tasks better served by Kling Video O1. [Routing] Choose Kling V3 Omni T2V when the user emphasizes Omni, consistency, multimodal quality, or complex text narratives. Prefer Kling V3 T2V as the default high-quality T2V baseline; prefer Kling V3 Turbo T2V when the user stresses speed, cost, or high-volume short-form output.
- kling/kling-v3-turbo-t2v: [Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical 720p/1080p short videos rather than maximum cinematic headroom. [Best For] Highly recommended for: rapid prototyping, social and ad iteration, batch short-form pipelines, and dialogue clips where turnaround time and unit cost matter most. [Limitations] Do NOT use this model if the user requires peak 4K cinematic fidelity, heavy multi-shot storyboard control, or maximum visual polish; use Kling V3 T2V or Kling V3 Omni T2V instead. Do NOT use it for image-conditioned animation; use Kling V3 Turbo I2V or Kling V3 I2V. [Routing] Choose Kling V3 Turbo T2V when the user says fast, quick, cheap, or high volume. Otherwise default to Kling V3 T2V for quality, or Kling V3 Omni T2V when consistency and Omni-class control are requested.
- kling/kling-v3-omni-i2v: [Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.
- kling/kling-v3-turbo-i2v: [Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.
- alibaba/wan2.7-image-pro-edit: [Core Function] Wan 2.7 Image Pro Edit is Alibaba’s flagship reasoning-enhanced image editing model. [Strengths] It supports interactive editing, character-consistent multi-image generation, and complex multi-reference modifications with deep reasoning. [Best For] Highly recommended for: professional image retouching, consistent character sheets, and complex structural edits. [Limitations] Do NOT use this model if you specifically need to use negative prompts to exclude elements during editing (use Qwen Image 2.0 Pro Edit instead). [Routing] Use this by default for high-end image editing and multi-reference consistent character generation.
- openai/gpt-image-2-edit: [Core Function] GPT Image 2 Edit is a high-resolution image-to-image editing model. [Strengths] It excels at making high-fidelity edits and style transformations to a single source image based on a text prompt, preserving details at up to 4K resolutions. [Best For] Highly recommended for: professional photo retouching, upscaling style transfers, and modifying high-resolution concept art. [Limitations] Do NOT use this model for multi-image merging (it only accepts one input image). Do NOT use if you need precise input fidelity control or transparent backgrounds. [Routing] Use this model by default when the user wants to edit a single image and prioritize output resolution/quality. If they need to merge multiple images or control the strictness of the edit (fidelity), use GPT Image 1.5 Edit.
- bytedance/seedream-5.0-lite-edit: [Core Function] Seedream 5.0 Lite Edit is a reasoning-enhanced, smart image editing model. [Strengths] It features superior cross-modal understanding and reasoning, allowing for highly accurate, interactive multi-turn image editing with real-time knowledge enhancement. [Best For] Highly recommended for: complex image editing tasks, structural modifications, and edits requiring deep semantic understanding. [Limitations] As a ‘Lite’ model, raw visual rendering might not match the 4.5 tier. [Routing] Use this model by default for complex, reasoning-based image editing tasks.
- alibaba/qwen-image-2.0-pro-edit: [Core Function] Qwen Image 2.0 Pro Edit is a highly controllable professional image editing model. [Strengths] It seamlessly unifies generation and editing, supporting NEGATIVE prompts during the editing process to strictly exclude elements. [Best For] Highly recommended for: precise image editing where specific elements must be removed or avoided. [Limitations] Does not feature the explicit ‘Thinking Mode’ reasoning of Wan 2.7. [Routing] Route to this model specifically when the user wants to edit an image AND provides a ‘negative prompt’ to exclude elements.
- minimax/minimax-image-01-i2i: [Core Function] MiniMax Image-01 I2I is an image-to-image editing and variation model. [Strengths] It excels at generating new images based on a text prompt while structurally referencing one or more input images. [Best For] Highly recommended for: style transfer, generating variations of existing artwork, and structurally guided image creation. [Limitations] Do NOT use this model if you want to generate video or if you do not have a reference image. [Routing] Use this model when the user provides a reference image and a text prompt to generate a new image.








