text-to-video
MiniMax H3 T2V

[Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.

AIモデルを探す

bytedance/seedance-2.0-t2v
bytedance/seedance-2.0-t2v
text-to-video

[Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.

alibaba/qwen-image-3.0-pro
alibaba/qwen-image-3.0-pro
text-to-image

[Core Function] Qwen Image 3.0 Pro is Alibaba's latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and long structured prompts for complex layouts. [Best For] Highly recommended for: photorealistic stills, marketing posters with readable text, detailed scene compositions, multi-panel layouts, product hero shots, and multi-variant creative exploration (n up to 6). [Limitations] Do NOT use this if the user needs native 4K output, thinking-mode reasoning, or image editing with reference images (use Qwen Image 3.0 Pro Edit for edits). Keep total pixels within 512*512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer this over Qwen Image 2.0 Pro for new Qwen Image text-to-image work. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Pro Edit instead.

kling/kling-v3-t2i
kling/kling-v3-t2i
text-to-image

[Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.

kling/kling-v3-t2v
kling/kling-v3-t2v
text-to-video

[Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.

alibaba/qwen-audio-3.0-tts-plus
alibaba/qwen-audio-3.0-tts-plus
text-to-speech

[Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanlingxin and longanlufeng (see the Qwen-Audio-TTS voice list). [Best For] Premium narration, brand voiceovers, multilingual product audio, and quality-sensitive batch TTS when Qwen-Audio voices are preferred. [Limitations] Do NOT mix Flash-only voices (e.g. longanhuan_v3.6) with Plus. [Routing] Choose Plus when speech quality is the priority. Choose qwen-audio-3.0-tts-flash when lower latency matters more.

xai/grok-imagine-image-quality
xai/grok-imagine-image-quality
text-to-image

[Core Function] Grok Imagine Image (Quality) is xAI's high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed concept art, marketing visuals, and any scenario where image quality is prioritized over generation speed. [Limitations] Do NOT use this model when latency is critical, as generation is slower than the standard model. [Routing] Use this model by default when the user emphasizes quality or detail. For faster, lighter generation use Grok Imagine Image (standard).

xai/grok-imagine-video
xai/grok-imagine-video
text-to-video

[Core Function] Grok Imagine Video is xAI's text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social clips, animated concepts, and dynamic scene generation from a description. [Limitations] Do NOT use this model when you have a starting image or reference subjects, or when you need resolutions above 720p or clips longer than 15 seconds; it is limited to 480p/720p and 15s. [Routing] Use this model when the user wants a video from text only. If a starting image is provided, route to the Image-to-Video model; for reference-driven character video, use Reference-to-Video.

minimax/hailuo-2.3-t2v
minimax/hailuo-2.3-t2v
text-to-video

[Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.

最新リリース

minimax/minimax-h3-v2v
minimax/minimax-h3-v2v
video-to-video

[Core Function] MiniMax H3 V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: motion remix, style transfer from video clips, keeping subject motion while changing the scene description, and multi-clip reference guidance. [Limitations] Do NOT use this model for text-only generation or first or last frame transitions. prompt, duration, resolution, and reference_videos are required. Do NOT send first_frame_image or last_frame_image on this endpoint. [Routing] Choose MiniMax H3 V2V when the user provides reference videos. For reference stills only, use MiniMax H3 I2V. For start or end frames, use MiniMax H3 FL2V. For text only, use MiniMax H3 T2V.

minimax/minimax-h3-i2v
minimax/minimax-h3-i2v
image-to-video

[Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions or when the primary input is a reference video. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 I2V when the user provides reference stills. For start or end frames, use MiniMax H3 FL2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

minimax/minimax-h3-fl2v
minimax/minimax-h3-fl2v
image-to-video

[Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 I2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.

minimax/minimax-h3-t2v
minimax/minimax-h3-t2v
text-to-video

[Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.

bytedance/seedance-2.5-v2v
bytedance/seedance-2.5-v2v
video-to-video

[Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 4-30 second output and mp4 or mov containers. [Best For] Highly recommended for: editing existing clips, extending motion from a base video, multi-reference restyling, and prompt-driven composition that cites Video n or Image n. [Limitations] Do NOT use this model for text-only or first-frame-only workflows. video_urls is required. Do NOT send first_frame_image or last_frame_image. Do NOT use this model when the user requires 1080p or 4k output. Do NOT use this model for audio-only input. [Routing] Prefer this model for Seedance 2.5 edit, extension, and video-reference jobs. Use Seedance 2.5 T2V for text-only and Seedance 2.5 I2V for image-first generation.

bytedance/seedance-2.5-i2v
bytedance/seedance-2.5-i2v
image-to-video

[Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 4-30 second duration, and mp4 or mov output. [Best For] Highly recommended for: animating a keyframe, first-to-last transitions, multi-image character consistency, and image-led storytelling on Seedance 2.5. [Limitations] Do NOT mix first_frame_image or last_frame_image with reference_images. Do NOT send video_urls on I2V. Do NOT use audio_urls alone. Do NOT use this model when the user requires 1080p or 4k output. [Routing] Prefer this model for Seedance 2.5 image-driven generation. Use Seedance 2.5 T2V for text-only requests and Seedance 2.5 V2V when reference video is required.

bytedance/seedance-2.5-t2v
bytedance/seedance-2.5-t2v
text-to-video

[Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-to-video storytelling, social clips beyond 15 seconds, and Seedance workflows that need mov output. [Limitations] Do NOT use this model if the user provides images, video, or audio as inputs. Do NOT use this model when the user requires 1080p or 4k output. [Routing] Prefer Seedance 2.5 T2V when the user needs more than 15 seconds of text-to-video. Use Seedance 2.0 T2V when 1080p or 4k is required. Use Seedance 2.5 I2V or V2V when media inputs are provided.

alibaba/wan3.0-v2v
alibaba/wan3.0-v2v
video-to-video

[Core Function] Wan 3.0 V2V is Alibaba Wan 3.0 reference-video generation that builds new video from one or more input videos. [Strengths] It supports up to 5 reference videos with optional reference images and audio for multimodal composition and prompt-referenced subjects. [Best For] Highly recommended for: video-to-video transformation, multi-subject scenes that cite video1/image1 in the prompt, and extending creative edits from existing clips. [Limitations] Do NOT use this model for text-only, file/link, or first-frame-only workflows. video_urls is required. [Routing] Prefer this model when the user supplies reference video. Use Wan 3.0 T2V for prompt/file/link and Wan 3.0 I2V for image-first generation.

シリーズ

CosyVoice

CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.

4 モデル
Fun

Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.

2 モデル
Gemini Omni

Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.

4 モデル
GPT Image

The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.

4 モデル
Grok Imagine

Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.

10 モデル
Grok Voice

Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.

2 モデル
Hailuo 02

MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.

3 モデル
Hailuo 2.3

MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.

3 モデル
HappyHorse

HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.

7 モデル
Kling V3

Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.

10 モデル
MAI Image

Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.

4 モデル
Nano Banana

Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.

8 モデル
PixVerse C1

PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.

4 モデル
PixVerse V6

PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.

5 モデル
Qwen Audio

Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.

2 モデル
Qwen Image

Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.

16 モデル
Seedance

ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.

18 モデル
Seedream

ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.

9 モデル
SkyReels

SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.

4 モデル
Veo 3

Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.

4 モデル
Veo 3.1

Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.

6 モデル
Vidu Q3

Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.

12 モデル
Wan

Alibaba's Wan (Wanx) series is a family of open-source multimodal foundation models developed by Alibaba Cloud that excels at high-quality text-to-video and text-to-image generation, featuring precise motion control, multilingual text rendering, and advanced instruction-following capabilities.

38 モデル

プロバイダー

Alibaba

Alibaba Cloud is a leading provider of advanced AI models, featuring the Qwen series (including Qwen-Image for multimodal vision-language tasks) and the Wan series for high-fidelity video generation.

70 モデル
Bytedance

ByteDance is a leading provider of advanced AI media models, featuring the Seedance series for high-fidelity multimodal video generation and the Seedream series for superior image creation and editing.

27 モデル
Google

Google is a leading provider of advanced AI media models, featuring Nano Banana and Imagen for high-fidelity image generation and editing, and Veo for scalable video synthesis.

25 モデル
Kling

Kuaishou is a leading provider of advanced AI media models, featuring the Kling series (including Video and Image) for high-fidelity multimodal video and image generation.

19 モデル
Microsoft

Microsoft is a global technology company founded by Bill Gates and Paul Allen in 1975, best known for its Windows operating system, Office productivity suite, Azure cloud platform, and its mission to "empower every person and every organization on the planet to achieve more."

5 モデル
MiniMax

MiniMax is a leading provider of advanced AI media models, featuring Hailuo for high-fidelity multimodal video generation and editing.

22 モデル
OpenAI

OpenAI is an AI research and deployment company founded in 2015, dedicated to developing safe and beneficial artificial general intelligence (AGI) that benefits all of humanity.

5 モデル
PixVerse

PixVerse is an AI-powered video generation platform that transforms text prompts and images into high-quality, creative short videos with features like real-time interaction, AI effects, and one-click storytelling.

13 モデル
Skywork

SkyWork is a versatile AI image generator that transforms text prompts into photorealistic, professional-grade visuals with lightning-fast speed, offering rich customization across artistic styles, colors, lighting, and composition in a seamless all-in-one creative workspace.

10 モデル
Vidu

Vidu is a cutting-edge AI video generation model that transforms text prompts, images, and reference videos into high-quality cinematic content. It supports multiple modes including text-to-video and image-to-video, delivering fast generation with strong motion consistency and 1080p resolution.

21 モデル
xAI

xAI, founded by Elon Musk in 2023, is an artificial intelligence company best known for its Grok chatbot — a witty, rebellious AI integrated into the X platform — with a stated mission to "understand the true nature of the universe."

12 モデル

カテゴリー

テキストから画像

Modellixでテキストから画像モデルとAPIを検索できます。

32 モデル
画像編集

Modellixで画像編集モデルとAPIを検索できます。

36 モデル
テキストから動画

Modellixでテキストから動画モデルとAPIを検索できます。

38 モデル
画像から動画

Modellixで画像から動画モデルとAPIを検索できます。

75 モデル
動画から動画

Modellixで動画から動画モデルとAPIを検索できます。

32 モデル
テキスト読み上げ

Modellixでテキスト読み上げモデルとAPIを検索できます。

9 モデル
音声認識

Modellixで音声認識モデルとAPIを検索できます。

5 モデル
音声変換

Modellixで音声変換モデルとAPIを検索できます。

2 モデル

コレクション

AI Animation Generator

The best animation generation Models.

16 モデル
AI Anime Generator

The best models for generating manga/anime images or videos.

14 モデル
AI Art Generator

AI models suitable for art design.

9 モデル
AI Avatar Generator

The best avatar-generating AI models.

9 モデル
AI Portrait Generator

The best portrait-generating AI models.

6 モデル
AI Style Transfer

This is a collection of the best models for style transfer, including image generation and video generation models.

14 モデル
AI Video Generator

It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.

30 モデル
Colorize Photo

The mainstream image-colorization AI models.

7 モデル
Digital Human

This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.

2 モデル
Face Swap

The mainstream face-swap models.

6 モデル
Image Upscale

The mainstream image upscalers.

6 モデル
Lip Sync

A curated collection of lip-sync AI models.

3 モデル
Logo Generator

The best LOGO generation design and creation AI models.

6 モデル
One Click

The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.

4 モデル
Photo Restoration

An AI models for restoring old photos.

7 モデル
Virtual Try-On

Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.

2 モデル
Voice Cloning

The best voice-cloning models.

2 モデル

主要なAIメディアモデルで開発

透明性の高い料金、Playgroundでの試用、本番ワークロード向けの統合APIで、画像・動画・音声モデルを利用できます。

検索を開始
必要なモデルが見つかりませんか?ご要望をお聞かせください。
動画生成145
画像生成68
音声生成11