google/imagen-4.0-generate-001

imagen-4.0-generate-001
ドキュメント
スキーマ

[Core Function] Imagen 4.0 is Google's flagship professional-grade text-to-image model designed for high-fidelity visual generation. [Strengths] It excels at industry-leading photorealism, exceptional typography/text rendering inside images, and strong prompt adherence, supporting up to 2K resolution and batch generation. [Best For] Highly recommended for general high-quality image generation, marketing assets (posters, product labels, menus, signage), and product photography mockups where readable text is required. [Limitations] Do NOT use this model if you need instant/real-time generation, native image editing (inpainting/outpainting), or negative prompts. It is strictly an image-only output model. [Routing] Use this model by default for high-quality, photorealistic text-to-image requests. If the user emphasizes speed or needs lower cost, route to Imagen 4.0 Fast. If the user demands the absolute highest detail and prompt precision, route to Imagen 4.0 Ultra.

$0.0322/img
text-to-image

入力

Image description text
Image aspect ratio
Output image resolution
Person generation policy
Number of images to generate per request

結果

結果はまだありません

モデルを実行すると、ここで出力をプレビューできます。

次へ:

README

"## What Imagen 4.0 Does
Imagen 4.0 is Google’s standard text-to-image model, delivering high-quality, photorealistic output from text prompts. It supports batch generation, person control, and resolution up to 2K — a strong all-round choice for realistic imagery.

Key Features

  • Photorealistic Quality: Produces high-quality, true-to-life images from text.
  • Up to 2K Resolution: Select 1K or 2K output via imageSize.
  • Batch Output: Generate 1–4 images per request via sampleCount.
  • Common Aspect Ratios: Supports 1:1, 3:4, 4:3, 9:16, and 16:9.
  • Person Generation Control: Use personGeneration (dont_allow, allow_adult, allow_all) to govern depiction of people.
  • Asynchronous Architecture: Uses a ““Create Task → Poll Result”” API workflow for reliable integration.

How to Use Imagen 4.0

  1. Write a specific, detailed text prompt describing the scene, subject, and style.
  2. Submit a POST request to the /google/imagen-4.0-generate-001 endpoint with your prompt and optional sampleCount, aspectRatio, personGeneration, and imageSize.
  3. Receive a task_id with a ““pending”” status in the response.
  4. Poll the provided GET endpoint to monitor the task.
  5. Once complete, retrieve and download the generated image(s).

Best Use Cases

  • Photorealistic Marketing Visuals: Create realistic product, lifestyle, and campaign imagery.
  • Editorial & Concept Imagery: Generate detailed scenes for articles and presentations.
  • Design Mockups: Visualize realistic environments and products for client work.
  • Batch Asset Creation: Produce coordinated image sets in a single request.

Tips for Better Results

  • Be Specific for Photorealism: Use detailed prompts describing lighting, materials, and composition for the best realistic output.
  • Use 2K for Final Renders: Set imageSize: 2K when you need maximum detail.
  • Batch with sampleCount: Generate up to 4 variations per request to pick the strongest result.
  • Let Aspect Ratio Guide Composition: Choose a ratio that supports your intended framing.
  • Step Up to Ultra When Needed: For the absolute highest detail and photorealism, use Imagen 4.0 Ultra."

属性

シリーズ

コレクション

パラメーター

名前 説明 必須 列挙値
prompt Image description text string はい -
aspectRatio Image aspect ratio string いいえ 1:1, 3:4, 4:3, 9:16, 16:9
imageSize Output image resolution string いいえ 1K, 2K
personGeneration Person generation policy string いいえ dont_allow, allow_adult, allow_all
sampleCount Number of images to generate per request integer いいえ -

料金

単位: $/img

料金
$0.0322/img

関連モデル

  • google/imagen-4.0-fast-generate-001: [Core Function] Imagen 4.0 Fast is a high-speed text-to-image model optimized for rapid visual generation. [Strengths] It offers incredible generation speed (approximately 2.7 seconds) and high cost-effectiveness while maintaining excellent text rendering inside images. [Best For] Highly recommended for rapid prototyping, quick design iteration, high-volume image generation, social media thumbnails, blog hero images, and UI mockups. [Limitations] Do NOT use this model if you require 2K resolution (it does not support the imageSize parameter) or the absolute highest cinematic fidelity and strict prompt alignment. Does not support negative prompts or native image editing. [Routing] Choose this model when the user emphasizes ‘fast’, ‘quick’, ‘rapid’, or needs high-volume, low-latency generation.
  • google/imagen-4.0-ultra-generate-001: [Core Function] Imagen 4.0 Ultra is Google’s premium text-to-image model optimized for the highest quality and instruction alignment. [Strengths] It delivers ultimate photorealistic detail, superior texture rendering, and extremely strict adherence to complex, multi-part prompt instructions, supporting up to 2K resolution. [Best For] Highly recommended for high-end commercial marketing assets, professional design, intricate artistic compositions, and scenarios where precise instruction-following is paramount. [Limitations] Do NOT use this model if you need instant/real-time generation, native image editing (inpainting/outpainting), or negative prompts. It is the most expensive tier in the Imagen family. [Routing] Choose this model when the user demands the absolute highest quality, ultimate detail, or has highly complex prompt instructions. Otherwise, default to standard Imagen 4.0 or Imagen 4.0 Fast.
  • openai/gpt-image-1.5: [Core Function] GPT Image 1.5 is a versatile text-to-image generation model. [Strengths] It balances solid visual performance with crucial utility features, notably its native support for generating images with transparent backgrounds. [Best For] Highly recommended for: creating UI icons, standalone logos, game assets, and any graphic design elements that require a transparent background. [Limitations] Do NOT use this model if you need 2K or 4K resolution. Its maximum supported resolution is 1536x1024. [Routing] Choose this model specifically when the user asks for ‘transparent background’, ‘no background’, or ‘PNG icon’. For standard, high-fidelity, or 4K image generation, use GPT Image 2 instead.
  • alibaba/wan2.7-image-pro: [Core Function] Wan 2.7 Image Pro is Alibaba’s flagship reasoning-enhanced image generation model. [Strengths] It features built-in chain-of-thought reasoning (Thinking Mode), exceptional prompt accuracy, native 12-language text rendering, and generates ultra-high-resolution 4K images. [Best For] Highly recommended for: print-ready large-format posters, complex logical prompts, and generating images containing specific text/typography. [Limitations] Do NOT use this model if you need to generate batch images rapidly (use Wan 2.7 Image instead) or if you specifically need negative prompts (use Qwen Image 2.0 Pro). [Routing] Use this model by default for high-end, 4K, or text-heavy image generation tasks.
  • alibaba/qwen-image-2.0-pro: [Core Function] Qwen Image 2.0 Pro is a professional-grade, highly controllable image generation model. [Strengths] It provides exquisite photorealism, professional infographics generation, supports NEGATIVE prompts, and can generate up to 6 image variants per API call. [Best For] Highly recommended for: workflows requiring strict negative prompt exclusion, batch generation (6 variants), and highly detailed infographics. [Limitations] Does not natively output 4K resolution (max is 2K). Does not have the explicit ‘Thinking Mode’ of Wan 2.7. [Routing] Route to this model specifically when the user provides a ‘negative prompt’ or asks for ‘batch generation of 6 images’.
  • bytedance/seedream-5.0-lite: [Core Function] Seedream 5.0 Lite is a smart, reasoning-enhanced image generation model with real-time web search capabilities. [Strengths] It excels at generating time-sensitive imagery, infographics, and content requiring deep world knowledge or online search, boasting superior prompt understanding and reasoning. [Best For] Highly recommended for: current-event posters, text-heavy designs, and concept art requiring complex logical reasoning. [Limitations] As a ‘Lite’ model, its absolute photorealistic aesthetic ceiling might be slightly lower than the specialized 4.5 model. [Routing] Use this model by default when the user needs real-time information, deep reasoning, or complex intent understanding in their image.
  • bytedance/seedream-5.0-pro: [Core Function] Seedream 5.0 Pro is ByteDance’s flagship professional-grade Text-to-Image (T2I) generation model. [Strengths] It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation consistency for professional scenarios. [Best For] Highly recommended for: professional design assets, high-fidelity photorealistic imagery, precisely controlled compositions, and brand or commercial visuals where quality matters most. [Limitations] Do NOT use this model for batch image generation or streaming output; it generates exactly one image per request and supports up to 2K resolution (no 3K/4K). [Routing] Choose this Pro model when the user emphasizes ultimate quality or precision. Choose Seedream 5.0 Lite when real-time web knowledge, batch generation, or 3K resolution is needed. To edit an existing image use Seedream 5.0 Pro Edit; to blend multiple reference images use Seedream 5.0 Pro Multi-Reference.
  • alibaba/qwen-image-3.0-pro: [Core Function] Qwen Image 3.0 Pro is Alibaba’s latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (widthheight), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and long structured prompts for complex layouts. [Best For] Highly recommended for: photorealistic stills, marketing posters with readable text, detailed scene compositions, multi-panel layouts, product hero shots, and multi-variant creative exploration (n up to 6). [Limitations] Do NOT use this if the user needs native 4K output, thinking-mode reasoning, or image editing with reference images (use Qwen Image 3.0 Pro Edit for edits). Keep total pixels within 512512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer this over Qwen Image 2.0 Pro for new Qwen Image text-to-image work. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Pro Edit instead.
  • kling/kling-v3-t2i: [Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.
  • openai/gpt-image-2: [Core Function] GPT Image 2 is a state-of-the-art text-to-image generation model. [Strengths] It excels at generating highly detailed, photorealistic images from text descriptions, with native support for ultra-high resolutions including 2K and 4K (up to 3840x2160). [Best For] Highly recommended for: cinematic landscapes, detailed character portraits, high-end commercial concept art, and any scenario requiring maximum resolution and visual fidelity. [Limitations] Do NOT use this model if you need a transparent background (e.g., for icons or UI assets), as it does not support the background: transparent parameter. [Routing] Use this model by default for all high-quality image generation requests. If the user explicitly asks for an image with a transparent background, route to GPT Image 1.5 instead.
  • xai/grok-imagine-image-quality: [Core Function] Grok Imagine Image (Quality) is xAI’s high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed concept art, marketing visuals, and any scenario where image quality is prioritized over generation speed. [Limitations] Do NOT use this model when latency is critical, as generation is slower than the standard model. [Routing] Use this model by default when the user emphasizes quality or detail. For faster, lighter generation use Grok Imagine Image (standard).

関連リソース