課金単位
to-imageUSD/imgで課金されます。「img」は画像枚数を表し、出力画像1枚単位で計算します。to-videoUSD/secで課金されます。「sec」は動画の長さを表し、1秒単位で計算します。to-audioUSD/M charsで課金されます。「M chars」は入力100万文字を表します。
to-imageUSD/imgで課金されます。「img」は画像枚数を表し、出力画像1枚単位で計算します。to-videoUSD/secで課金されます。「sec」は動画の長さを表し、1秒単位で計算します。to-audioUSD/M charsで課金されます。「M chars」は入力100万文字を表します。一部のモデルでは入力パラメーターに応じて料金が変わります。たとえば、特定の動画生成モデルでは、 1080p 動画の生成料金が 720p 動画より高く設定されています。
利用の多いモデルの公開料金です。モデルを選択すると、パラメーター別の料金を確認できます。
[Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism (use the Imagen 4 series instead). [Routing] Route to this model for 'fast', 'creative', or 'stylized' high-volume requests.
[Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.
[Core Function] GPT Image 2 is a state-of-the-art text-to-image generation model. [Strengths] It excels at generating highly detailed, photorealistic images from text descriptions, with native support for ultra-high resolutions including 2K and 4K (up to 3840x2160). [Best For] Highly recommended for: cinematic landscapes, detailed character portraits, high-end commercial concept art, and any scenario requiring maximum resolution and visual fidelity. [Limitations] Do NOT use this model if you need a transparent background (e.g., for icons or UI assets), as it does not support the `background: transparent` parameter. [Routing] Use this model by default for all high-quality image generation requests. If the user explicitly asks for an image with a transparent background, route to GPT Image 1.5 instead.
[Core Function] Wan 2.7 T2V is Alibaba's flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: high-quality commercial video generation, professional storytelling, and dynamic cinematic sequences. [Limitations] Do NOT use this model if the user specifically requests the streamlined 'HappyHorse' workflow. [Routing] Use this model by default for high-end text-to-video requests on the Alibaba platform.
[Core Function] Seedream 5.0 Lite is a smart, reasoning-enhanced image generation model with real-time web search capabilities. [Strengths] It excels at generating time-sensitive imagery, infographics, and content requiring deep world knowledge or online search, boasting superior prompt understanding and reasoning. [Best For] Highly recommended for: current-event posters, text-heavy designs, and concept art requiring complex logical reasoning. [Limitations] As a 'Lite' model, its absolute photorealistic aesthetic ceiling might be slightly lower than the specialized 4.5 model. [Routing] Use this model by default when the user needs real-time information, deep reasoning, or complex intent understanding in their image.
[Core Function] Seedream 5.0 Pro is ByteDance's flagship professional-grade Text-to-Image (T2I) generation model. [Strengths] It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation consistency for professional scenarios. [Best For] Highly recommended for: professional design assets, high-fidelity photorealistic imagery, precisely controlled compositions, and brand or commercial visuals where quality matters most. [Limitations] Do NOT use this model for batch image generation or streaming output; it generates exactly one image per request and supports up to 2K resolution (no 3K/4K). [Routing] Choose this Pro model when the user emphasizes ultimate quality or precision. Choose Seedream 5.0 Lite when real-time web knowledge, batch generation, or 3K resolution is needed. To edit an existing image use Seedream 5.0 Pro Edit; to blend multiple reference images use Seedream 5.0 Pro Multi-Reference.
| モデル | タイプ | 料金 | $10で生成できる上限 |
|---|---|---|---|
| vidu/vidu-q2-ns-i2v | 画像から動画 | 動画 43本(各5秒) | |
| vidu/vidu-new-ns-i2v | 画像から動画 | 動画 18本(各5秒) | |
| alibaba/happyhorse-1.1-t2v | テキストから動画 | 動画 16本(各5秒) | |
| alibaba/happyhorse-1.0-t2v | テキストから動画 | 動画 16本(各5秒) | |
| alibaba/happyhorse-1.1-r2v | 画像から動画 | 動画 16本(各5秒) | |
| alibaba/happyhorse-1.1-i2v | 画像から動画 | 動画 16本(各5秒) | |
| alibaba/happyhorse-1.0-r2v | 画像から動画 | 動画 16本(各5秒) | |
| alibaba/happyhorse-1.0-i2v | 画像から動画 | 動画 16本(各5秒) | |
| alibaba/happyhorse-1.0-video-edit | 動画から動画 | 動画 16本(各5秒) | |
| alibaba/qwen-image-3.0-pro-edit | 画像から画像 | 画像 217枚 | |
| alibaba/qwen-image-3.0-pro | テキストから画像 | 画像 217枚 | |
| kling/kling-video-effects | 画像から動画 | 動画 3本(各5秒) | |
| kling/kling-avatar | 画像から動画 | 動画 51本(各5秒) | |
| kling/kling-video-o1 | 動画から動画 | 動画 23本(各5秒) | |
| kling/kling-v3-omni-video | 動画から動画 | 動画 23本(各5秒) | |
| kling/kling-video-o1-i2v | 画像から動画 | 動画 34本(各5秒) | |
| kling/kling-v3-turbo-i2v | 画像から動画 | 動画 25本(各5秒) | |
| kling/kling-v3-omni-i2v | 画像から動画 | 動画 34本(各5秒) | |
| kling/kling-v3-i2v | 画像から動画 | 動画 34本(各5秒) | |
| kling/kling-video-o1-t2v | テキストから動画 | 動画 34本(各5秒) |
ユースケース別にまとめた各コレクションで、モデルの最低料金を確認できます。
It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.
Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.
This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.
This is a collection of the best models for style transfer, including image generation and video generation models.
A curated collection of lip-sync AI models.
The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.
The best animation generation Models.
The best models for generating manga/anime images or videos.
AI models suitable for art design.
The best avatar-generating AI models.
The best portrait-generating AI models.
The mainstream image-colorization AI models.
The mainstream face-swap models.
The mainstream image upscalers.
The best LOGO generation design and creation AI models.
An AI models for restoring old photos.
The best voice-cloning models.
Lite、Fast、Standardなどのバリエーションを、シリーズ単位で確認できます。
Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.
ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.
Alibaba's Wan 2.7 series is a comprehensive open-weight AI suite for image generation/editing and video creation, featuring thinking mode reasoning, first/last frame control, up to 4K images and 1080p videos, native audio sync, and exceptional text rendering accuracy.
Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.
ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.
The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.
CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.
Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.
Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.
Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.
Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.
MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.
MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.
HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.
Google Imagen is Google's premier text-to-image diffusion model, excelling in photorealistic, high-resolution image generation from textual prompts with unmatched detail, creativity, and adherence to complex descriptions.
Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.
Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.
MiniMax Speech is a series of advanced text-to-speech (TTS) models that deliver ultra-low latency, highly natural and expressive speech synthesis, with support for zero-shot voice cloning and multilingual capabilities across variants like Turbo and HD.
PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.
PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.
Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.
Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.
SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.
Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.
Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.
Alibaba's Wan 2.6 is a powerful open-source AI video generation model that creates cinematic 1080p multi-shot videos with native audio-visual synchronization, supporting text-to-video, image-to-video, and professional storytelling workflows.
Modellix上の各プロバイダーの最低料金を比較できます。
Google is a leading provider of advanced AI media models, featuring Nano Banana and Imagen for high-fidelity image generation and editing, and Veo for scalable video synthesis.
ByteDance is a leading provider of advanced AI media models, featuring the Seedance series for high-fidelity multimodal video generation and the Seedream series for superior image creation and editing.
Alibaba Cloud is a leading provider of advanced AI models, featuring the Qwen series (including Qwen-Image for multimodal vision-language tasks) and the Wan series for high-fidelity video generation.
OpenAI is an AI research and deployment company founded in 2015, dedicated to developing safe and beneficial artificial general intelligence (AGI) that benefits all of humanity.
Kuaishou is a leading provider of advanced AI media models, featuring the Kling series (including Video and Image) for high-fidelity multimodal video and image generation.
MiniMax is a leading provider of advanced AI media models, featuring Hailuo for high-fidelity multimodal video generation and editing.
生成タイプ別の最低料金です。必要な機能に対応する課金単位を確認できます。
ありません。Modellixは透明性の高い従量課金を採用しています。公開されているモデル料金と実際の利用量に基づいてお支払いいただきます。
プロバイダーの原価やモデルの提供状況が変わった場合、モデル料金が変更されることがあります。最新料金は常に料金ページとモデル詳細ページに表示されます。
大規模または長期の利用については、アカウント単位のチャージ割引やクレジット割引を提供する場合があります。モデル料金は公開料金表に基づきます。
はい。コンソールで「Billing > Orders」を開き、対象注文の「Actions」列にある「Invoice」をクリックするとダウンロードできます。請求情報のカスタマイズや個別要件については、support@modellix.aiまでお問い合わせください。
モデル名と料金で検索されるブログやドキュメントの関連コンテンツを確認できます。
生成中です。このページを閉じないでください...
多くの開発者やクリエイターが利用しています。今すぐ、すべてのモデルに対応する統合APIをお試しください。