AIモデルAPI料金

標準化されたパラメーターを1つのAI APIで提供。隠れたコストはありません。すべてのモデルの公開料金により、予算を正確に管理できます。

料金の仕組み

課金単位

  • to-imageUSD/imgで課金されます。「img」は画像枚数を表し、出力画像1枚単位で計算します。
  • to-videoUSD/secで課金されます。「sec」は動画の長さを表し、1秒単位で計算します。
  • to-audioUSD/M charsで課金されます。「M chars」は入力100万文字を表します。

パラメーターによる料金差

一部のモデルでは入力パラメーターに応じて料金が変わります。たとえば、特定の動画生成モデルでは、 1080p 動画の生成料金が 720p 動画より高く設定されています。

コレクション別料金

ユースケース別にまとめた各コレクションで、モデルの最低料金を確認できます。

AI Video Generator

It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.

最低 $0.0575/sec

Virtual Try-On

Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.

最低 $0.09/img

Digital Human

This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.

最低 $0.092/sec

AI Style Transfer

This is a collection of the best models for style transfer, including image generation and video generation models.

最低 $0.004/img

Lip Sync

A curated collection of lip-sync AI models.

最低 $0.0115/sec

One Click

The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.

最低 $0.0552/sec

AI Animation Generator

The best animation generation Models.

最低 $0.0414/sec

AI Anime Generator

The best models for generating manga/anime images or videos.

最低 $0.004/img

AI Art Generator

AI models suitable for art design.

最低 $0.0041/img

AI Avatar Generator

The best avatar-generating AI models.

最低 $0.0138/img

AI Portrait Generator

The best portrait-generating AI models.

最低 $0.0041/img

Colorize Photo

The mainstream image-colorization AI models.

最低 $0.0117/img

Face Swap

The mainstream face-swap models.

最低 $0.004/img

Image Upscale

The mainstream image upscalers.

最低 $0.0117/img

Logo Generator

The best LOGO generation design and creation AI models.

最低 $0.0138/img

Photo Restoration

An AI models for restoring old photos.

最低 $0.0117/img

Voice Cloning

The best voice-cloning models.

最低 $8.0637/M chars

シリーズ別料金

Lite、Fast、Standardなどのバリエーションを、シリーズ単位で確認できます。

Nano Banana

Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.

最低 $0.0274/img

Seedance

ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.

最低 $0.0124/sec

Wan 2.7

Alibaba's Wan 2.7 series is a comprehensive open-weight AI suite for image generation/editing and video creation, featuring thinking mode reasoning, first/last frame control, up to 4K images and 1080p videos, native audio sync, and exceptional text rendering accuracy.

最低 $0.0185/img

Veo 3.1

Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.

最低 $0.0403/sec

Seedream

ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.

最低 $0.031/img

GPT Image

The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.

最低 $0.0041/img

CosyVoice

CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.

最低 $8.0637/M chars

Fun

Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.

最低 $0.002/sec

Gemini Omni

Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.

最低 $0.0805/sec

Grok Imagine

Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.

最低 $0.023/img

Grok Voice

Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.

最低 $0.0001/sec

Hailuo 02

MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.

最低 $0.02/sec

Hailuo 2.3

MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.

最低 $0.04/sec

HappyHorse

HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.

最低 $0.1207/sec

Imagen

Google Imagen is Google's premier text-to-image diffusion model, excelling in photorealistic, high-resolution image generation from textual prompts with unmatched detail, creativity, and adherence to complex descriptions.

最低 $0.0161/img

Kling V3

Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.

最低 $0.0386/img

MAI Image

Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.

最低 $0.0388/img

MiniMax Speech

MiniMax Speech is a series of advanced text-to-speech (TTS) models that deliver ultra-low latency, highly natural and expressive speech synthesis, with support for zero-shot voice cloning and multilingual capabilities across variants like Turbo and HD.

料金情報なし

PixVerse C1

PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.

最低 $0.069/sec

PixVerse V6

PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.

最低 $0.0575/sec

Qwen Audio

Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.

最低 $9.4702/M chars

Qwen Image

Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.

最低 $0.0172/img

SkyReels

SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.

最低 $0.0805/sec

Veo 3

Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.

最低 $0.115/sec

Vidu Q3

Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.

最低 $0.0184/sec

Wan 2.6

Alibaba's Wan 2.6 is a powerful open-source AI video generation model that creates cinematic 1080p multi-shot videos with native audio-visual synchronization, supporting text-to-video, image-to-video, and professional storytelling workflows.

最低 $0.0152/sec

カテゴリー別料金

生成タイプ別の最低料金です。必要な機能に対応する課金単位を確認できます。

テキストから画像

テキストプロンプトから静止画を生成します。出力画像1枚単位で課金されます。

最低 $0.004/img

画像から画像

既存画像を編集、再構成、補正します。出力画像1枚単位で課金されます。

最低 $0.004/img

テキストから動画

テキストプロンプトから動画クリップを生成します。出力動画1秒単位で課金されます。

最低 $0.0116/sec

画像から動画

元画像またはキーフレームに動きを加えます。出力動画1秒単位で課金されます。

最低 $0.0008/sec

動画から動画

既存動画のスタイル変更、延長、変換を行います。1秒単位で課金されます。

最低 $0.0115/sec

テキスト読み上げ

テキストを自然な音声に変換し、音声ワークフローに利用できます。

最低 $8.0637/M chars

よくある質問

隠れた手数料はありますか?

ありません。Modellixは透明性の高い従量課金を採用しています。公開されているモデル料金と実際の利用量に基づいてお支払いいただきます。

料金はどのくらいの頻度で変更されますか?

プロバイダーの原価やモデルの提供状況が変わった場合、モデル料金が変更されることがあります。最新料金は常に料金ページとモデル詳細ページに表示されます。

ボリュームディスカウントの仕組みは?

大規模または長期の利用については、アカウント単位のチャージ割引やクレジット割引を提供する場合があります。モデル料金は公開料金表に基づきます。

カスタム請求書を発行できますか?

はい。コンソールで「Billing > Orders」を開き、対象注文の「Actions」列にある「Invoice」をクリックするとダウンロードできます。請求情報のカスタマイズや個別要件については、support@modellix.aiまでお問い合わせください。

料金・導入ガイド

モデル名と料金で検索されるブログやドキュメントの関連コンテンツを確認できます。

すべての記事を見る

生成中です。このページを閉じないでください...

Modellixで開発を始めませんか?

多くの開発者やクリエイターが利用しています。今すぐ、すべてのモデルに対応する統合APIをお試しください。

今すぐ生成を始める