AI Model API Pricing

One AI API, flattened parameters. No hidden costs. Public AI API pricing for every model helps you budget with total precision.

How Pricing Works

Pricing Units

  • to-imageCharged by USD/img, where "img" is the number of images, counted per image.
  • to-videoCharged by USD/sec, where "sec" is the video duration, counted per second.
  • to-audioCharged by USD/M chars, where "M chars" means one million input characters.

Parameter Variations

Some models use different prices based on input parameters. For example, certain video generation models have higher prices for generating 1080p videos compared to 720p videos.

Pricing by Collection

Use-case hubs with starting prices across the models in each collection.

AI Video Generator

It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.

from $0.043/sec

Virtual Try-On

Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.

from $0.07/img

Digital Human

This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.

from $0.005/sec

AI Style Transfer

This is a collection of the best models for style transfer, including image generation and video generation models.

from $0.022/img

Lip Sync

A curated collection of lip-sync AI models.

from $0.012/sec

One Click

The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.

from $0.06/sec

AI Animation Generator

The best animation generation Models.

from $0.045/sec

AI Anime Generator

The best models for generating manga/anime images or videos.

from $0.0224/img

AI Art Generator

AI models suitable for art design.

from $0.0054/img

AI Avatar Generator

The best avatar-generating AI models.

from $0.005/sec

AI Portrait Generator

The best portrait-generating AI models.

from $0.0054/img

Colorize Photo

The mainstream image-colorization AI models.

from $0.0135/img

Face Swap

The mainstream face-swap models.

from $0.0224/img

Image Upscale

The mainstream image upscalers.

from $0.0224/img

Logo Generator

The best LOGO generation design and creation AI models.

from $0.0117/img

Photo Restoration

An AI models for restoring old photos.

from $0.0135/img

Voice Cloning

The best voice-cloning models.

from $13/M chars

Pricing by Series

Family hubs that group variants (Lite / Fast / Standard) under one series page.

Nano Banana

Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.

from $0.0306/img

Seedance

ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.

from $0.021/sec

Veo 3.1

Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.

from $0.045/sec

Seedream

ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.

from $0.035/img

GPT Image

The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.

from $0.0054/img

CosyVoice

CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.

from $13/M chars

Fun

Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.

from $0.0001/sec

Gemini Omni

Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.

from $0.1/sec

Grok Imagine

Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.

from $0.024/img

Grok Voice

Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.

from $0.0001/sec

Hailuo 02

MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.

from $0.0153/sec

Hailuo 2.3

MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.

from $0.0288/sec

HappyHorse

HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.

from $0.14/sec

Imagen

Google Imagen is Google's premier text-to-image diffusion model, excelling in photorealistic, high-resolution image generation from textual prompts with unmatched detail, creativity, and adherence to complex descriptions.

Pricing unavailable

Kling V3

Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.

from $0.0224/img

MAI Image

Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.

from $0.0195/img

MiniMax Speech

MiniMax Speech is a series of advanced text-to-speech (TTS) models that deliver ultra-low latency, highly natural and expressive speech synthesis, with support for zero-shot voice cloning and multilingual capabilities across variants like Turbo and HD.

Pricing unavailable

PixVerse C1

PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.

from $0.06/sec

PixVerse V6

PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.

from $0.05/sec

Qwen Audio

Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.

from $15/M chars

Qwen Image

Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.

from $0.027/img

SkyReels

SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.

from $0.084/sec

Veo 3

Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.

Pricing unavailable

Vidu Q3

Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.

from $0.035/sec

Wan

Alibaba's Wan (Wanx) series is a family of open-source multimodal foundation models developed by Alibaba Cloud that excels at high-quality text-to-video and text-to-image generation, featuring precise motion control, multilingual text rendering, and advanced instruction-following capabilities.

from $0.027/img

Pricing By Category

Starting prices by generation type — match billing unit to the capability you need.

Text to Image

Generate still images from text prompts. Billed per output image.

from $0.0054/img

Image to Image

Edit, remix, and refine existing images. Billed per output image.

from $0.0135/img

Text to Video

Generate video clips from text prompts. Billed per second of output.

from $0.0338/sec

Image to Video

Animate a source image or keyframe into motion. Billed per second of output.

from $0.005/sec

Video to Video

Restyle, extend, or transform existing video. Billed per second.

from $0.012/sec

Text to Speech

Turn written text into natural speech and audio for voice workflows.

from $13/M chars

Frequently Asked Questions

Are there any hidden fees?

No. Modellix uses transparent usage-based pricing. You pay based on the public model price and your actual usage.

How often do prices change?

Model prices may change when provider costs or model availability changes. The latest price is always shown on the pricing page and model detail pages.

How do volume discounts work?

For high-volume or long-term usage, Modellix may offer account-level top-up or credit discounts. Model prices remain based on the public pricing table.

Can I get a custom invoice?

Yes. Go to the Console, open Billing > Orders, and click Invoice in the Actions column of the relevant order to download it yourself. For custom billing details or special requirements, contact support@modellix.ai.

Pricing & Integration Guides

Close the loop with blog/docs content that already ranks for model + pricing queries.

Browse All Articles

Generating, please do not close this page...

Ready to build with Modellix?

Join thousands of developers and creators. Start using the unified API for all models today.

Start Creating Now