Modellixで本番環境に対応した28種類の画像編集 AIモデルを検索。機能を比較し、Playgroundで試して、1つの統合APIから導入できます。

[Core Function] GPT Image 2.5 Flare Edit is a speed-oriented image editing model. [Strengths] Supports up to 16 input images, optional mask-based local edits, resolution tiers up to 4K, and transparent backgrounds. [Best For] Product image changes, masked object replacements, and edits guided by multiple reference images. [Limitations] Do NOT use this for editing more than 16 input images or selecting arbitrary pixel dimensions. Exact input fidelity control is not exposed. [Routing] Choose this variant when speed is the priority; use GPT Image 2.5 Sunburst Edit when image quality is the priority.

[Core Function] GPT Image 2.5 Sunburst Edit is a quality-oriented image editing model. [Strengths] Supports up to 16 input images, optional mask-based local edits, resolution tiers up to 4K, and transparent backgrounds. [Best For] Product image changes, masked object replacements, and edits guided by multiple reference images. [Limitations] Do NOT use this for editing more than 16 input images or selecting arbitrary pixel dimensions. Exact input fidelity control is not exposed. [Routing] Choose this variant when image quality is the priority; use GPT Image 2.5 Flare Edit when speed is the priority.

[Core Function] MAI Image 2.6 Flash Edit is the faster, lower-cost variant of MAI Image 2.6 image editing. [Strengths] It applies prompt-guided edits to a single source image with lower latency than MAI Image 2.6 Edit. [Best For] Highly recommended for: high-throughput edit APIs and production retouching pipelines. [Limitations] Do NOT use this if the caller provides more than one source image, a data URI, or a URL that is not publicly reachable over HTTP or HTTPS. It accepts exactly one JPEG or PNG image URL, and output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when edit speed or cost matters more than maximum 2.6 quality. For the highest fidelity, route to MAI Image 2.6 Edit.

[Core Function] MAI Image 2.6 Edit is Microsoft's latest prompt-guided image editing model in the MAI Image family. [Strengths] It applies targeted edits to a single source image with the same quality gains as MAI Image 2.6 generation. [Best For] Highly recommended for: object edits, layout changes, text cleanup, and iterative photorealistic retouching. [Limitations] Do NOT use this if the caller provides more than one source image, a data URI, or a URL that is not publicly reachable over HTTP or HTTPS. It accepts exactly one JPEG or PNG image URL, and output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when the latest MAI edit quality is required. For faster or lower-cost edits, route to MAI Image 2.6 Flash Edit.

[Core Function] Grok Imagine Image 2.0 Edit edits one to three source images from a text prompt. [Strengths] It supports quality (low or medium), 1k or 2k resolution, and the same aspect ratios as Grok Imagine Image 2.0. [Best For] Highly recommended for: restyling, combining up to 3 references, and iterative refinement. [Limitations] Requires at least one source image; at most 3 images per request. This model does not generate from text alone. [Routing] For text-to-image generation, use Grok Imagine Image 2.0.

[Core Function] MAI Image 2.5 Pro Edit is Microsoft's highest-fidelity image editing model in the MAI 2.5 family. [Strengths] It excels at applying premium, prompt-guided edits and transformations to a single source image. [Best For] Highly recommended for: hero asset retouching, high-fidelity restyling, and detailed prompt-driven edits. [Limitations] Do NOT provide more than one source image or non-JPEG/PNG inputs; it accepts exactly one JPEG or PNG image, and output is always PNG. [Routing] Use this model when maximum edit quality is required. For faster or lower-cost edits, route to MAI Image 2.5 Edit or MAI Image 2.5 Flash Edit.

[Core Function] Qwen Image 3.0 Edit is the standard image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite (direct mode), and 1-6 outputs while preserving subject identity. [Best For] Background replacement, outfit or style changes, multi-image fusion, and iterative retouching when Pro-tier quality is not required. [Limitations] Do NOT use this for pure text-to-image with no reference images (use Qwen Image 3.0 instead). Keep output total pixels within 512*512 to 2048*2048. prompt_extend_mode only supports direct (agent is T2I-only). [Routing] Route here when the user provides reference image(s) and wants balanced Qwen 3.0 edit quality. Prefer Qwen Image 3.0 Pro Edit for higher quality edits.

[Core Function] Qwen Image 3.0 Pro Edit is an image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite, 1-6 outputs while preserving subject identity, and long structured edit instructions. [Best For] Highly recommended for: background replacement, outfit or style changes, multi-image fusion, identity-preserving portrait edits, and iterative creative retouching. [Limitations] Do NOT use this for pure text-to-image with no reference images (use Qwen Image 3.0 Pro instead), native 4K output, or thinking-mode reasoning. Keep output total pixels within 512*512 to 2048*2048; input images should follow supported formats and size guidance. Do NOT combine very long prompts with multiple reference images and a long negative_prompt if the request may exceed the model input capacity (about 4.5k tokens total across text and images). [Routing] Route here when the user provides reference image(s) and wants Qwen 3.0 edit quality. For text-only generation without images, use Qwen Image 3.0 Pro.

[Core Function] Seedream 5.0 Pro Multi-Reference is a professional-grade multi-reference image generation (I2I) model that creates a single image from 2-10 reference images plus a text prompt. [Strengths] It excels at reference consistency, preserving characters, styles, and objects across multiple input images while following complex blending instructions with professional-grade quality. [Best For] Highly recommended for: keeping character or style consistency across references, combining subjects from different images into one scene, placing products into reference scenes, and IP-consistent content creation. [Limitations] Do NOT use this model with fewer than 2 or more than 10 reference images, and do NOT use it for batch generation or streaming; it outputs exactly one image per request. [Routing] For single-image editing use Seedream 5.0 Pro Edit; for text-only generation use Seedream 5.0 Pro; choose Seedream 5.0 Lite Edit when up to 14 reference images or batch outputs are needed.

[Core Function] Seedream 5.0 Pro Edit is a professional-grade single-image editing (I2I) model. [Strengths] It supports interactive precise editing: edit locations can be specified via coordinates, selection boxes, or arrows described in the prompt, with strong element-level control and subject consistency. [Best For] Highly recommended for: precise local retouching, adding/removing/replacing objects at exact positions, style transfer of a single photo, and professional post-editing workflows. [Limitations] Do NOT use this model for text-to-image generation (an input image is required) or for blending multiple reference images; it accepts exactly one input image and outputs exactly one image (no batch or streaming). [Routing] For 2-10 reference images use Seedream 5.0 Pro Multi-Reference; for pure text-to-image use Seedream 5.0 Pro; choose Seedream 5.0 Lite Edit when batch outputs or more than 10 input images are required.

[Core Function] Nano Banana 2 Lite Edit (Gemini 3.1 Flash Lite Image) is the lightweight, cost-efficient image editing model of the Nano Banana 2 family; it transforms one or more input images per a text instruction. [Strengths] It performs fast, low-cost instruction-based editing across up to 14 input images. [Best For] Highly recommended for: high-volume edits, quick style transforms, and batch background or attribute changes where cost and throughput matter. [Limitations] Do NOT use this model when you need the highest edit fidelity or richest detail; use Nano Banana 2 Edit or Nano Banana Pro Edit instead. [Routing] Choose the Lite variant for cost- and throughput-sensitive edits; step up to Nano Banana 2 Edit for higher quality.

[Core Function] Grok Imagine Image Edit is xAI's standard image editing model. [Strengths] It excels at quickly applying prompt-guided edits and style changes to one or more source images. [Best For] Highly recommended for: fast restyling, quick variations, and lightweight image edits. [Limitations] Do NOT use this model when maximum edit fidelity is required; the Quality variant preserves more detail. A maximum of 3 source images is supported. [Routing] Choose this model for fast edits. When the user demands maximum fidelity, route to Grok Imagine Image Edit (Quality).

[Core Function] Grok Imagine Image Edit (Quality) is xAI's high-fidelity image editing model. [Strengths] It excels at applying detailed, prompt-guided edits and style transformations to one or more source images while preserving fine detail. [Best For] Highly recommended for: high-quality restyling, detailed inpainting-style edits, and combining up to 3 source images. [Limitations] Do NOT use this model when latency is critical, as it is slower than the standard edit variant; a maximum of 3 source images is supported. [Routing] Use this model by default for quality-sensitive edits. For faster edits, route to Grok Imagine Image Edit (standard).

[Core Function] MAI Image 2.5 Flash Edit is Microsoft's fast, cost-efficient image editing model. [Strengths] It excels at quickly applying prompt-guided edits to a single source image. [Best For] Highly recommended for: rapid edits, batch processing, and cost-sensitive workloads. [Limitations] Do NOT provide more than one source image or non-JPEG/PNG inputs; it accepts exactly one JPEG or PNG image, output is always PNG, and fidelity is lower than MAI Image 2.5 Edit. [Routing] Choose this model when speed or cost matters more than maximum fidelity. For the highest quality, use MAI Image 2.5 Edit.

[Core Function] MAI Image 2.5 Edit is Microsoft's flagship image editing model. [Strengths] It excels at applying high-quality, prompt-guided edits and transformations to a single source image. [Best For] Highly recommended for: restyling, object/scene modification, and detailed prompt-driven edits. [Limitations] Do NOT provide more than one source image or non-JPEG/PNG inputs; it accepts exactly one JPEG or PNG image, and output is always PNG. [Routing] Use this model by default for quality-sensitive edits. For faster, cheaper edits, route to MAI Image 2.5 Flash Edit.

[Core Function] GPT Image 2 Edit is a high-resolution image-to-image editing model. [Strengths] It excels at making high-fidelity edits and style transformations to a single source image based on a text prompt, preserving details at up to 4K resolutions. [Best For] Highly recommended for: professional photo retouching, upscaling style transfers, and modifying high-resolution concept art. [Limitations] Do NOT use this model for multi-image merging (it only accepts one input image). Do NOT use if you need precise input fidelity control or transparent backgrounds. [Routing] Use this model by default when the user wants to edit a single image and prioritize output resolution/quality. If they need to merge multiple images or control the strictness of the edit (fidelity), use GPT Image 1.5 Edit.

[Core Function] Kling V3 Omni Image is a unified multimodal image generation endpoint (POST /images/omni-image). [Strengths] Multi-image reference, up to 4K, and optional series generation via result_type/series_amount. Use <<<image_N>>> placeholders in prompt. [Best For] Character consistency, fusing multiple reference images, and comic/storyboard series. [Limitations] Element library IDs are not exposed. Default aspect_ratio is 1:1 when omitted. [Routing] Prefer for multi-image fusion or series; use kling-v3-t2i/i2i for standard single-shot generation.

[Core Function] Kling V3 I2I is the flagship image-to-image editing model (POST /images/generations, model_name=kling-v3 with image). [Strengths] High-quality style transfer and editing up to 2K. [Best For] Single-reference image editing. [Limitations] Do NOT send negative_prompt when image is present (officially unsupported). No image_fidelity / image_reference on V3. For multi-image fusion or series, use kling-v3-omni-image. [Routing] Default for standard image-to-image.

[Core Function] Wan 2.7 Image Edit is a fast, reasoning-enhanced image editing model. [Strengths] Provides the robust editing capabilities of the Wan 2.7 architecture with faster turnaround times. [Best For] Highly recommended for: standard image modifications and style transfers. [Limitations] Do NOT use if you need absolute maximum fidelity or negative prompt support. [Routing] Use for standard, fast image editing tasks.

[Core Function] Wan 2.7 Image Pro Edit is Alibaba's flagship reasoning-enhanced image editing model. [Strengths] It supports interactive editing, character-consistent multi-image generation, and complex multi-reference modifications with deep reasoning. [Best For] Highly recommended for: professional image retouching, consistent character sheets, and complex structural edits. [Limitations] Do NOT use this model if you specifically need to use negative prompts to exclude elements during editing (use Qwen Image 2.0 Pro Edit instead). [Routing] Use this by default for high-end image editing and multi-reference consistent character generation.

[Core Function] Seedream 5.0 Lite Edit is a reasoning-enhanced, smart image editing model. [Strengths] It features superior cross-modal understanding and reasoning, allowing for highly accurate, interactive multi-turn image editing with real-time knowledge enhancement. [Best For] Highly recommended for: complex image editing tasks, structural modifications, and edits requiring deep semantic understanding. [Limitations] As a 'Lite' model, raw visual rendering might not match the 4.5 tier. [Routing] Use this model by default for complex, reasoning-based image editing tasks.

[Core Function] Nano Banana 2 Edit is a high-speed image editing model. [Strengths] It rapidly modifies existing images or extracts image frames from videos based on text prompts. [Best For] Highly recommended for: rapid style transfer, quick image modifications, and fast creative edits. [Limitations] Do NOT use this model for meticulous photorealistic retouching. [Routing] Use this model by default for fast, creative image editing tasks.

[Core Function] Nano Banana Pro Edit is a high-capability creative image editing model. [Strengths] It provides high-quality creative edits, background replacements, and style transformations. [Best For] Highly recommended for: detailed creative modifications and complex style transfers. [Limitations] Do NOT use for strict photorealistic restoration. [Routing] Use this model for detailed, high-quality creative edits.

[Core Function] Nano Banana Edit is the original fast image editing model. [Strengths] Fast basic edits. [Limitations] Superseded by Nano Banana 2 Edit. [Routing] Default to Nano Banana 2 Edit unless specifically requested.

[Core Function] Kolors Virtual Try-On V1.5 is a specialized AI fashion model. [Strengths] It highly accurately applies garments (including Top+Bottom combinations) onto a person's image, preserving fabric texture and draping. [Best For] Highly recommended for: e-commerce virtual fitting rooms and fashion visualization. [Limitations] Do NOT use this model for general image editing. It is strictly for clothing try-on. [Routing] Use this model whenever the user asks to 'try on' clothes or apply a garment to a person.

[Core Function] Kolors Virtual Try-On V1 is a legacy AI fashion and virtual try-on model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy e-commerce integrations that have not yet migrated to the newer try-on pipeline. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kolors Virtual Try-On V1.5.

[Core Function] Kling Image Expansion is an outpainting model. [Strengths] It intelligently extends the borders of an image (horizontal, vertical, or asymmetric) while matching the original style and context. [Best For] Highly recommended for: changing aspect ratios, extending landscapes, and filling out cropped photos. [Limitations] Do NOT use this model for inpainting or style transfer. [Routing] Use this specifically when the user asks to 'expand', 'extend', or 'uncrop' an image.

[Core Function] Kling Image O1 is a reasoning-enhanced multimodal image model. [Strengths] It performs deep reasoning over prompts and references to handle complex logic, spatial relationships, and intricate multi-image combinations. [Best For] Highly recommended for: complex scenes requiring strict logical or spatial accuracy. [Limitations] Element library IDs and series generation are not exposed. Default aspect_ratio is 1:1 when omitted. Do NOT use for simple artistic generation where V3 is faster and more stylistic. [Routing] Route to this model when the prompt involves complex physical logic or strict spatial reasoning.