计费单位
to-image按 USD/img 计费,其中 img 表示图片数量,按张计费。to-video按 USD/sec 计费,其中 sec 表示视频时长,按秒计费。to-audio按 USD/M chars 计费,其中 M chars 表示一百万个输入字符。
to-image按 USD/img 计费,其中 img 表示图片数量,按张计费。to-video按 USD/sec 计费,其中 sec 表示视频时长,按秒计费。to-audio按 USD/M chars 计费,其中 M chars 表示一百万个输入字符。部分模型会根据输入参数采用不同价格。例如,某些视频生成模型生成 1080p 视频时,价格会高于 720p 视频。
查看热门模型的公开价格,进入详情页可了解不同参数下的定价。
[Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism; prefer Nano Banana Pro for higher-fidelity creative output. [Routing] Route to this model for 'fast', 'creative', or 'stylized' high-volume requests.
[Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.
[Core Function] GPT Image 2 is a state-of-the-art text-to-image generation model. [Strengths] It excels at generating highly detailed, photorealistic images from text descriptions, with native support for ultra-high resolutions including 2K and 4K (up to 3840x2160). [Best For] Highly recommended for: cinematic landscapes, detailed character portraits, high-end commercial concept art, and any scenario requiring maximum resolution and visual fidelity. [Limitations] Do NOT use this model if you need a transparent background (e.g., for icons or UI assets), as it does not support the `background: transparent` parameter. [Routing] Use this model by default for all high-quality image generation requests. If the user explicitly asks for an image with a transparent background, route to GPT Image 1.5 instead.
[Core Function] Seedream 5.0 Lite is a smart, reasoning-enhanced image generation model with real-time web search capabilities. [Strengths] It excels at generating time-sensitive imagery, infographics, and content requiring deep world knowledge or online search, boasting superior prompt understanding and reasoning. [Best For] Highly recommended for: current-event posters, text-heavy designs, and concept art requiring complex logical reasoning. [Limitations] As a 'Lite' model, its absolute photorealistic aesthetic ceiling might be slightly lower than the specialized 4.5 model. [Routing] Use this model by default when the user needs real-time information, deep reasoning, or complex intent understanding in their image.
[Core Function] MAI Image 2.6 is Microsoft's latest text-to-image generation model in the MAI Image family. [Strengths] It improves text rendering, portraits, 3D imagery, and commercial photorealistic output compared with MAI Image 2.5. [Best For] Highly recommended for: marketing visuals, product hero shots, portraits, and prompts that need accurate on-image text. [Limitations] Do NOT use this if the requested width or height is below 768, or if width x height exceeds 1,048,576 pixels. Output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when the latest MAI quality is required. For lower latency or cost, route to MAI Image 2.6 Flash.
[Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.
| 模型 | 类型 | 价格 | 最多生成($10) |
|---|---|---|---|
| microsoft/mai-image-2.6-flash-edit | 图生图 | 454 张图 | |
| microsoft/mai-image-2.6-edit | 图生图 | 212 张图 | |
| microsoft/mai-image-2.6-flash | 文生图 | 512 张图 | |
| microsoft/mai-image-2.6 | 文生图 | 257 张图 | |
| bytedance/seedance-2.5-v2v | 视频生视频 | 32 个 5 秒视频 | |
| bytedance/seedance-2.5-i2v | 图生视频 | 19 个 5 秒视频 | |
| bytedance/seedance-2.5-t2v | 文生视频 | 19 个 5 秒视频 | |
| bytedance/seedream-5.0-pro-multi-reference | 图生图 | 123 张图 | |
| bytedance/seedream-5.0-pro-edit | 图生图 | 123 张图 | |
| bytedance/seedream-5.0-pro | 文生图 | 123 张图 | |
| bytedance/seedance-2.0-mini-v2v | 视频生视频 | 95 个 5 秒视频 | |
| bytedance/seedance-2.0-mini-i2v | 图生视频 | 50 个 5 秒视频 | |
| bytedance/seedance-2.0-mini-t2v | 文生视频 | 50 个 5 秒视频 | |
| bytedance/seedance-2.0-fast-t2v | 文生视频 | 33 个 5 秒视频 | |
| bytedance/seedance-2.0-t2v | 文生视频 | 28 个 5 秒视频 | |
| bytedance/seedream-5.0-lite-edit | 图生图 | 285 张图 | |
| bytedance/seedream-5.0-lite | 文生图 | 285 张图 | |
| bytedance/seedance-2.0-fast-v2v | 视频生视频 | 60 个 5 秒视频 | |
| bytedance/seedance-2.0-v2v | 视频生视频 | 46 个 5 秒视频 | |
| bytedance/seedance-2.0-fast-i2v | 图生视频 | 33 个 5 秒视频 |
按使用场景浏览模型集合,并查看每个集合中的最低起步价。
It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.
Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.
This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.
This is a collection of the best models for style transfer, including image generation and video generation models.
A curated collection of lip-sync AI models.
The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.
The best animation generation Models.
The best models for generating manga/anime images or videos.
AI models suitable for art design.
The best avatar-generating AI models.
The best portrait-generating AI models.
The mainstream image-colorization AI models.
The mainstream face-swap models.
The mainstream image upscalers.
The best LOGO generation design and creation AI models.
An AI models for restoring old photos.
The best voice-cloning models.
在同一系列页中比较 Lite、Fast、Standard 等不同模型版本。
Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.
ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.
Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.
ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.
The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.
CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.
Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.
Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.
Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.
Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.
MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.
MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.
HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.
Google Imagen is Google's premier text-to-image diffusion model, excelling in photorealistic, high-resolution image generation from textual prompts with unmatched detail, creativity, and adherence to complex descriptions.
Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.
Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.
MiniMax Speech is a series of advanced text-to-speech (TTS) models that deliver ultra-low latency, highly natural and expressive speech synthesis, with support for zero-shot voice cloning and multilingual capabilities across variants like Turbo and HD.
PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.
PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.
Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.
Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.
SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.
Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.
Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.
Alibaba's Wan (Wanx) series is a family of open-source multimodal foundation models developed by Alibaba Cloud that excels at high-quality text-to-video and text-to-image generation, featuring precise motion control, multilingual text rendering, and advanced instruction-following capabilities.
比较 Modellix 上各模型供应商产品目录的最低起步价。
Google is a leading provider of advanced AI media models, featuring Nano Banana and Imagen for high-fidelity image generation and editing, and Veo for scalable video synthesis.
ByteDance is a leading provider of advanced AI media models, featuring the Seedance series for high-fidelity multimodal video generation and the Seedream series for superior image creation and editing.
Alibaba Cloud is a leading provider of advanced AI models, featuring the Qwen series (including Qwen-Image for multimodal vision-language tasks) and the Wan series for high-fidelity video generation.
OpenAI is an AI research and deployment company founded in 2015, dedicated to developing safe and beneficial artificial general intelligence (AGI) that benefits all of humanity.
Kuaishou is a leading provider of advanced AI media models, featuring the Kling series (including Video and Image) for high-fidelity multimodal video and image generation.
MiniMax is a leading provider of advanced AI media models, featuring Hailuo for high-fidelity multimodal video generation and editing.
没有隐藏费用,也没有订阅费。你只需要根据公开模型价格和实际使用量付费。
模型价格可能会随供应商成本或模型可用性变化而调整。最新价格以 Pricing 页面和模型详情页展示为准。
高用量或长期使用客户可联系我们咨询账户充值或额度折扣。模型单价仍以公开价格表为准。
可以。前往控制台,进入「计费 > 订单」,点击对应订单「操作」列的 Invoice 即可自助下载发票。如需定制抬头或特殊开票需求,请邮件联系 support@modellix.ai。
通过已有博客和文档,深入了解模型接入方式与定价细节。
正在生成中,请勿关闭此页面...
加入开发者和创作者队伍,从今天开始使用覆盖所有模型的统一 API。