Attributes
- Provider: openai
- Category: Text to Image
Series
Collections
Parameters
| Name | Description | Type | Required | Enums |
|---|---|---|---|---|
| prompt | Image description, from 1 to 32,000 Unicode characters. | string | Yes | - |
| aspect_ratio | Output aspect ratio. Can be set independently of resolution; omitting it uses 1:1. | string | No | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 |
| background | Background mode: auto lets the model choose, opaque requests an opaque background, and transparent requests a transparent background with PNG output. | string | No | auto, opaque, transparent |
| quality | Output quality level. Defaults to low. | string | No | low, medium, high, xhigh, max |
| resolution | Output resolution tier. Can be set independently of aspect_ratio; omitting it uses 1K. Pixel dimensions depend on both parameters: 1:1 + 1K gives 1024x1024, 1:1 + 4K gives 2880x2880, and 16:9 + 4K gives 3840x2160. | string | No | 1K, 2K, 4K |
Pricing
Unit: $/img
| Dimension | Pricing |
|---|---|
| quality: low / resolution: 1K | 0.0053 |
| quality: low / resolution: 2K | 0.1072 |
| quality: low / resolution: 4K | 0.0178 |
| quality: medium / resolution: 1K | 0.0124 |
| quality: medium / resolution: 2K | 0.0241 |
| quality: medium / resolution: 4K | 0.0401 |
| quality: high / resolution: 1K | 0.0474 |
| quality: high / resolution: 2K | 0.0963 |
| quality: high / resolution: 4K | 0.1601 |
| quality: xhigh / resolution: 1K | 0.0843 |
| quality: xhigh / resolution: 2K | 0.1710 |
| quality: xhigh / resolution: 4K | 0.2846 |
| quality: max / resolution: 1K | 0.1890 |
| quality: max / resolution: 2K | 0.3780 |
| quality: max / resolution: 4K | 0.6390 |
Related Models
- openai/gpt-image-2: [Core Function] GPT Image 2 is a state-of-the-art text-to-image generation model. [Strengths] It excels at generating highly detailed, photorealistic images from text descriptions, with native support for ultra-high resolutions including 2K and 4K (up to 3840x2160). [Best For] Highly recommended for: cinematic landscapes, detailed character portraits, high-end commercial concept art, and any scenario requiring maximum resolution and visual fidelity. [Limitations] Do NOT use this model if you need a transparent background (e.g., for icons or UI assets), as it does not support the
background: transparentparameter. [Routing] Use this model by default for all high-quality image generation requests. If the user explicitly asks for an image with a transparent background, route to GPT Image 1.5 instead. - openai/gpt-image-2-edit: [Core Function] GPT Image 2 Edit is a high-resolution image-to-image editing model. [Strengths] It excels at making high-fidelity edits and style transformations to a single source image based on a text prompt, preserving details at up to 4K resolutions. [Best For] Highly recommended for: professional photo retouching, upscaling style transfers, and modifying high-resolution concept art. [Limitations] Do NOT use this model for multi-image merging (it only accepts one input image). Do NOT use if you need precise input fidelity control or transparent backgrounds. [Routing] Use this model by default when the user wants to edit a single image and prioritize output resolution/quality. If they need to merge multiple images or control the strictness of the edit (fidelity), use GPT Image 1.5 Edit.
- openai/gpt-image-2.5-sunburst: [Core Function] GPT Image 2.5 Sunburst is a quality-oriented text-to-image generation model. [Strengths] Supports ten aspect ratios, resolution tiers up to 4K, and transparent backgrounds. [Best For] Landscape illustrations, product posters, and transparent icons. [Limitations] Do NOT use this for editing an existing image or selecting arbitrary pixel dimensions. [Routing] Choose this variant when image quality is the priority; use GPT Image 2.5 Flare when speed is the priority. Use GPT Image 2.5 Sunburst Edit when input images are required.
- openai/gpt-image-2.5-flare-edit: [Core Function] GPT Image 2.5 Flare Edit is a speed-oriented image editing model. [Strengths] Supports up to 16 input images, optional mask-based local edits, resolution tiers up to 4K, and transparent backgrounds. [Best For] Product image changes, masked object replacements, and edits guided by multiple reference images. [Limitations] Do NOT use this for editing more than 16 input images or selecting arbitrary pixel dimensions. Exact input fidelity control is not exposed. [Routing] Choose this variant when speed is the priority; use GPT Image 2.5 Sunburst Edit when image quality is the priority.
- openai/gpt-image-2.5-sunburst-edit: [Core Function] GPT Image 2.5 Sunburst Edit is a quality-oriented image editing model. [Strengths] Supports up to 16 input images, optional mask-based local edits, resolution tiers up to 4K, and transparent backgrounds. [Best For] Product image changes, masked object replacements, and edits guided by multiple reference images. [Limitations] Do NOT use this for editing more than 16 input images or selecting arbitrary pixel dimensions. Exact input fidelity control is not exposed. [Routing] Choose this variant when image quality is the priority; use GPT Image 2.5 Flare Edit when speed is the priority.
- google/nano-banana-2: [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism; prefer Nano Banana Pro for higher-fidelity creative output. [Routing] Route to this model for ‘fast’, ‘creative’, or ‘stylized’ high-volume requests.
- google/nano-banana-pro: [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [Limitations] Do NOT use this model when you need the absolute highest photorealism or detail; step up within the Nano Banana family (Nano Banana 2 or Nano Banana Pro Edit workflows) as needed. [Routing] Use this model for high-quality, creative, non-photorealistic requests.
- kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.
- xai/grok-imagine-image-quality: [Core Function] Grok Imagine Image (Quality) is xAI’s high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed concept art, marketing visuals, and any scenario where image quality is prioritized over generation speed. [Limitations] Do NOT use this model when latency is critical, as generation is slower than the standard model. [Routing] Use this model by default when the user emphasizes quality or detail. For faster, lighter generation use Grok Imagine Image (standard).
- microsoft/mai-image-2.5: [Core Function] MAI Image 2.5 is Microsoft’s flagship text-to-image generation model. [Strengths] It excels at producing high-quality, detailed images from a text prompt with precise control over output dimensions. [Best For] Highly recommended for: concept art, marketing visuals, and high-fidelity image generation. [Limitations] Do NOT request dimensions below 768 on any side or a width x height product above 1,048,576 pixels (e.g. beyond 1024x1024); output is always PNG. [Routing] Use this model by default for quality-sensitive generation. For faster, cheaper generation, route to MAI Image 2.5 Flash.
- bytedance/seedream-5.0-pro: [Core Function] Seedream 5.0 Pro is ByteDance’s flagship professional-grade Text-to-Image (T2I) generation model. [Strengths] It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation consistency for professional scenarios. [Best For] Highly recommended for: professional design assets, high-fidelity photorealistic imagery, precisely controlled compositions, and brand or commercial visuals where quality matters most. [Limitations] Do NOT use this model for batch image generation or streaming output; it generates exactly one image per request and supports up to 2K resolution (no 3K/4K). [Routing] Choose this Pro model when the user emphasizes ultimate quality or precision. Choose Seedream 5.0 Lite when real-time web knowledge, batch generation, or 3K resolution is needed. To edit an existing image use Seedream 5.0 Pro Edit; to blend multiple reference images use Seedream 5.0 Pro Multi-Reference.
- alibaba/qwen-image-3.0-pro: [Core Function] Qwen Image 3.0 Pro is Alibaba’s latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (widthheight), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and long structured prompts for complex layouts. [Best For] Highly recommended for: photorealistic stills, marketing posters with readable text, detailed scene compositions, multi-panel layouts, product hero shots, and multi-variant creative exploration (n up to 6). [Limitations] Do NOT use this if the user needs native 4K output, thinking-mode reasoning, or image editing with reference images (use Qwen Image 3.0 Pro Edit for edits). Keep total pixels within 512512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer this over Qwen Image 2.0 Pro for new Qwen Image text-to-image work. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Pro Edit instead.
- microsoft/mai-image-2.6: [Core Function] MAI Image 2.6 is Microsoft’s latest text-to-image generation model in the MAI Image family. [Strengths] It improves text rendering, portraits, 3D imagery, and commercial photorealistic output compared with MAI Image 2.5. [Best For] Highly recommended for: marketing visuals, product hero shots, portraits, and prompts that need accurate on-image text. [Limitations] Do NOT use this if the requested width or height is below 768, or if width x height exceeds 1,048,576 pixels. Output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when the latest MAI quality is required. For lower latency or cost, route to MAI Image 2.6 Flash.
- microsoft/mai-image-2.6-flash: [Core Function] MAI Image 2.6 Flash is the faster, lower-cost variant of MAI Image 2.6 text-to-image generation. [Strengths] It targets similar quality to MAI Image 2.6 with lower latency for high-throughput workloads. [Best For] Highly recommended for: production pipelines, batch generation, and latency-sensitive image APIs. [Limitations] Do NOT use this if the requested width or height is below 768, or if width x height exceeds 1,048,576 pixels. Output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when speed or cost matters more than maximum 2.6 quality. For the highest fidelity, route to MAI Image 2.6.
- alibaba/wan2.7-image-pro: [Core Function] Wan 2.7 Image Pro is Alibaba’s flagship reasoning-enhanced image generation model. [Strengths] It features built-in chain-of-thought reasoning (Thinking Mode), exceptional prompt accuracy, native 12-language text rendering, and generates ultra-high-resolution 4K images. [Best For] Highly recommended for: print-ready large-format posters, complex logical prompts, and generating images containing specific text/typography. [Limitations] Do NOT use this model if you need to generate batch images rapidly (use Wan 2.7 Image instead) or if you specifically need negative prompts (use Qwen Image 2.0 Pro). [Routing] Use this model by default for high-end, 4K, or text-heavy image generation tasks.
- bytedance/seedream-5.0-lite: [Core Function] Seedream 5.0 Lite is a smart, reasoning-enhanced image generation model with real-time web search capabilities. [Strengths] It excels at generating time-sensitive imagery, infographics, and content requiring deep world knowledge or online search, boasting superior prompt understanding and reasoning. [Best For] Highly recommended for: current-event posters, text-heavy designs, and concept art requiring complex logical reasoning. [Limitations] As a ‘Lite’ model, its absolute photorealistic aesthetic ceiling might be slightly lower than the specialized 4.5 model. [Routing] Use this model by default when the user needs real-time information, deep reasoning, or complex intent understanding in their image.







