AI 模型 API 定价

一个 AI API,统一参数,无隐藏费用。所有模型价格公开透明,助你精准规划每一分预算。

定价如何计算

计费单位

  • to-image按 USD/img 计费,其中 img 表示图片数量,按张计费。
  • to-video按 USD/sec 计费,其中 sec 表示视频时长,按秒计费。
  • to-audio按 USD/M chars 计费,其中 M chars 表示一百万个输入字符。

参数差异

部分模型会根据输入参数采用不同价格。例如,某些视频生成模型生成 1080p 视频时,价格会高于 720p 视频。

按集合查看定价

按使用场景浏览模型集合,并查看每个集合中的最低起步价。

AI Video Generator

It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.

起价 $0.043/sec

Virtual Try-On

Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.

起价 $0.07/img

Digital Human

This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.

起价 $0.005/sec

AI Style Transfer

This is a collection of the best models for style transfer, including image generation and video generation models.

起价 $0.022/img

Lip Sync

A curated collection of lip-sync AI models.

起价 $0.012/sec

One Click

The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.

起价 $0.06/sec

AI Animation Generator

The best animation generation Models.

起价 $0.045/sec

AI Anime Generator

The best models for generating manga/anime images or videos.

起价 $0.0224/img

AI Art Generator

AI models suitable for art design.

起价 $0.0054/img

AI Avatar Generator

The best avatar-generating AI models.

起价 $0.005/sec

AI Portrait Generator

The best portrait-generating AI models.

起价 $0.0054/img

Colorize Photo

The mainstream image-colorization AI models.

起价 $0.0135/img

Face Swap

The mainstream face-swap models.

起价 $0.0224/img

Image Upscale

The mainstream image upscalers.

起价 $0.0224/img

Logo Generator

The best LOGO generation design and creation AI models.

起价 $0.0117/img

Photo Restoration

An AI models for restoring old photos.

起价 $0.0135/img

Voice Cloning

The best voice-cloning models.

起价 $13/M chars

按系列查看定价

在同一系列页中比较 Lite、Fast、Standard 等不同模型版本。

Nano Banana

Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.

起价 $0.0306/img

Seedance

ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.

起价 $0.021/sec

Veo 3.1

Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.

起价 $0.045/sec

Seedream

ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.

起价 $0.035/img

GPT Image

The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.

起价 $0.0054/img

CosyVoice

CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.

起价 $13/M chars

Fun

Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.

起价 $0.0001/sec

Gemini Omni

Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.

起价 $0.1/sec

Grok Imagine

Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.

起价 $0.024/img

Grok Voice

Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.

起价 $0.0001/sec

Hailuo 02

MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.

起价 $0.0153/sec

Hailuo 2.3

MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.

起价 $0.0288/sec

HappyHorse

HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.

起价 $0.14/sec

Imagen

Google Imagen is Google's premier text-to-image diffusion model, excelling in photorealistic, high-resolution image generation from textual prompts with unmatched detail, creativity, and adherence to complex descriptions.

价格暂不可用

Kling V3

Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.

起价 $0.0224/img

MAI Image

Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.

起价 $0.0195/img

MiniMax Speech

MiniMax Speech is a series of advanced text-to-speech (TTS) models that deliver ultra-low latency, highly natural and expressive speech synthesis, with support for zero-shot voice cloning and multilingual capabilities across variants like Turbo and HD.

价格暂不可用

PixVerse C1

PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.

起价 $0.06/sec

PixVerse V6

PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.

起价 $0.05/sec

Qwen Audio

Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.

起价 $15/M chars

Qwen Image

Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.

起价 $0.027/img

SkyReels

SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.

起价 $0.084/sec

Veo 3

Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.

价格暂不可用

Vidu Q3

Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.

起价 $0.035/sec

Wan

Alibaba's Wan (Wanx) series is a family of open-source multimodal foundation models developed by Alibaba Cloud that excels at high-quality text-to-video and text-to-image generation, featuring precise motion control, multilingual text rendering, and advanced instruction-following capabilities.

起价 $0.027/img

按分类查看定价

按生成类型查看起步价,根据所需能力匹配对应计费单位。

文生图

根据文本提示生成静态图片,按输出图片张数计费。

起价 $0.0054/img

图生图

编辑、重绘和优化已有图片,按输出图片张数计费。

起价 $0.0135/img

文生视频

根据文本提示生成视频片段,按输出视频时长计费。

起价 $0.0338/sec

图生视频

将原始图片或关键帧生成动态视频,按输出视频时长计费。

起价 $0.005/sec

视频生视频

对已有视频进行风格转换、延展或重制,按视频时长计费。

起价 $0.012/sec

文生语音

将书面文本转换为自然语音,适用于配音和语音生成。

起价 $13/M chars

常见问题

是否有隐藏费用?

没有隐藏费用,也没有订阅费。你只需要根据公开模型价格和实际使用量付费。

价格多久会调整一次?

模型价格可能会随供应商成本或模型可用性变化而调整。最新价格以 Pricing 页面和模型详情页展示为准。

高用量折扣如何计算?

高用量或长期使用客户可联系我们咨询账户充值或额度折扣。模型单价仍以公开价格表为准。

是否可以开具定制发票?

可以。前往控制台,进入「计费 > 订单」,点击对应订单「操作」列的 Invoice 即可自助下载发票。如需定制抬头或特殊开票需求,请邮件联系 support@modellix.ai。

定价与接入指南

通过已有博客和文档,深入了解模型接入方式与定价细节。

浏览全部文章

正在生成中,请勿关闭此页面...

准备好用 Modellix 构建了吗?

加入开发者和创作者队伍,从今天开始使用覆盖所有模型的统一 API。

立即开始创作