
[Core Function] Seedream 5.0 Pro is ByteDance's flagship professional-grade Text-to-Image (T2I) generation model. [Strengths] It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation consistency for professional scenarios. [Best For] Highly recommended for: professional design assets, high-fidelity photorealistic imagery, precisely controlled compositions, and brand or commercial visuals where quality matters most. [Limitations] Do NOT use this model for batch image generation or streaming output; it generates exactly one image per request and supports up to 2K resolution (no 3K/4K). [Routing] Choose this Pro model when the user emphasizes ultimate quality or precision. Choose Seedream 5.0 Lite when real-time web knowledge, batch generation, or 3K resolution is needed. To edit an existing image use Seedream 5.0 Pro Edit; to blend multiple reference images use Seedream 5.0 Pro Multi-Reference.

[Core Function] Qwen Image 3.0 Pro is Alibaba's latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and long structured prompts for complex layouts. [Best For] Highly recommended for: photorealistic stills, marketing posters with readable text, detailed scene compositions, multi-panel layouts, product hero shots, and multi-variant creative exploration (n up to 6). [Limitations] Do NOT use this if the user needs native 4K output, thinking-mode reasoning, or image editing with reference images (use Qwen Image 3.0 Pro Edit for edits). Keep total pixels within 512*512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer this over Qwen Image 2.0 Pro for new Qwen Image text-to-image work. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Pro Edit instead.

[Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.

[Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.

[Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanlingxin and longanlufeng (see the Qwen-Audio-TTS voice list). [Best For] Premium narration, brand voiceovers, multilingual product audio, and quality-sensitive batch TTS when Qwen-Audio voices are preferred. [Limitations] Do NOT mix Flash-only voices (e.g. longanhuan_v3.6) with Plus. [Routing] Choose Plus when speech quality is the priority. Choose qwen-audio-3.0-tts-flash when lower latency matters more.

[Core Function] Grok Imagine Image (Quality) is xAI's high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed concept art, marketing visuals, and any scenario where image quality is prioritized over generation speed. [Limitations] Do NOT use this model when latency is critical, as generation is slower than the standard model. [Routing] Use this model by default when the user emphasizes quality or detail. For faster, lighter generation use Grok Imagine Image (standard).

[Core Function] Wan 2.7 T2V is Alibaba's flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: high-quality commercial video generation, professional storytelling, and dynamic cinematic sequences. [Limitations] Do NOT use this model if the user specifically requests the streamlined 'HappyHorse' workflow. [Routing] Use this model by default for high-end text-to-video requests on the Alibaba platform.

[Core Function] Grok Imagine Video is xAI's text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social clips, animated concepts, and dynamic scene generation from a description. [Limitations] Do NOT use this model when you have a starting image or reference subjects, or when you need resolutions above 720p or clips longer than 15 seconds; it is limited to 480p/720p and 15s. [Routing] Use this model when the user wants a video from text only. If a starting image is provided, route to the Image-to-Video model; for reference-driven character video, use Reference-to-Video.

[Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.

[Core Function] Vidu I2V New NS is an Image-to-Video model that animates a single first-frame image into an audio-visual synchronized, multi-shot narrative clip via the Vidu template API (vidu_i2v_new_ns). [Strengths] It excels at newer visual quality with multi-shot storytelling and longer durations from 2 to 15 seconds at 720p or 1080p (default 1080p). [Best For] Highly recommended for: cinematic storyboard-style clips from one still, brand films needing multi-shot continuity, longer narrative shorts up to 15s, and high-resolution audio-synced animation. [Limitations] Do NOT use this model if you need 360p or durations outside 2-15 seconds; it accepts only one input image and does not expose seed or aspect_ratio. [Routing] Choose Vidu I2V New NS for the latest multi-shot narrative I2V quality and longer clips. If the user needs 360p previews or only 3-10s with the q2-ns-duration template, route to Vidu Q2-NS instead.

[Core Function] Vidu Q2-NS is an Image-to-Video model that animates a single first-frame image into an audio-visual synchronized clip via the Vidu template API (q2-ns-duration). [Strengths] It excels at producing sound-synced motion from one still image with flexible durations from 3 to 10 seconds and resolutions including 360p for lower-cost previews. [Best For] Highly recommended for: short social clips with ambient audio, product motion from a hero still, cost-sensitive audio-visual drafts at 360p, and general first-frame animation with negative-prompt control. [Limitations] Do NOT use this model for multi-shot narrative storytelling or when you need durations outside 3-10 seconds; it accepts only one input image and does not expose seed or aspect_ratio. [Routing] Choose Vidu Q2-NS for audio-synced I2V with 360p option and 3-10s duration. If the user needs longer clips up to 15s, multi-shot narrative, or defaults to 1080p cinematic quality, route to Vidu I2V New NS instead.

[Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.

[Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical 720p/1080p short videos rather than maximum cinematic headroom. [Best For] Highly recommended for: rapid prototyping, social and ad iteration, batch short-form pipelines, and dialogue clips where turnaround time and unit cost matter most. [Limitations] Do NOT use this model if the user requires peak 4K cinematic fidelity, heavy multi-shot storyboard control, or maximum visual polish; use Kling V3 T2V or Kling V3 Omni T2V instead. Do NOT use it for image-conditioned animation; use Kling V3 Turbo I2V or Kling V3 I2V. [Routing] Choose Kling V3 Turbo T2V when the user says fast, quick, cheap, or high volume. Otherwise default to Kling V3 T2V for quality, or Kling V3 Omni T2V when consistency and Omni-class control are requested.

[Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.

[Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and better adherence when scenes demand coherent characters or multi-beat storytelling from text alone. [Best For] Highly recommended for: narrative T2V with recurring subjects, dialogue-aware scenes, brand or product continuity across beats, and premium short films where consistency matters more than raw throughput. [Limitations] Do NOT use this model if the user only needs the cheapest or fastest clip; prefer Kling V3 Turbo T2V. Do NOT use it when the workflow is image-first or needs multi-image references; use Kling V3 Omni I2V or Kling V3 I2V instead. Do NOT use it for deep physics-reasoning specialty tasks better served by Kling Video O1. [Routing] Choose Kling V3 Omni T2V when the user emphasizes Omni, consistency, multimodal quality, or complex text narratives. Prefer Kling V3 T2V as the default high-quality T2V baseline; prefer Kling V3 Turbo T2V when the user stresses speed, cost, or high-volume short-form output.

[Core Function] Kling Video O1 I2V is the image-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced generation from 1-7 reference images, 720p/1080p, duration 3-10s (single image only 5 or 10). [Best For] Complex physical motion grounded in reference frames. [Limitations] No native audio; no aspect_ratio (follows first frame); duration capped at 10s. [Routing] Prefer this for O1-quality I2V; use Kling Video O1 V2V when a source video is required.

[Core Function] Kling Video O1 T2V is the text-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced prompt planning with 3-10s duration and 720p/1080p output. [Best For] Complex physical interactions and logically demanding scenes from text alone. [Limitations] No native audio; duration capped at 10s; no multi_shot. [Routing] Prefer this for O1-quality T2V; use Kling Video O1 V2V when a source video is required.

CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.
4 モデル
Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.
2 モデル
Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.
4 モデル
The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.
4 モデル
Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.
10 モデル
Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.
2 モデル
MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.
3 モデル
MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.
3 モデル
HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.
7 モデル
Google Imagen is Google's premier text-to-image diffusion model, excelling in photorealistic, high-resolution image generation from textual prompts with unmatched detail, creativity, and adherence to complex descriptions.
3 モデル
Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.
10 モデル
Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.
4 モデル
Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.
8 モデル
PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.
4 モデル
PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.
5 モデル
Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.
2 モデル
Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.
14 モデル
ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.
15 モデル
ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.
9 モデル
SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.
4 モデル
Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.
4 モデル
Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.
6 モデル
Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.
12 モデル
Alibaba's Wan 2.6 is a powerful open-source AI video generation model that creates cinematic 1080p multi-shot videos with native audio-visual synchronization, supporting text-to-video, image-to-video, and professional storytelling workflows.
7 モデル
Alibaba's Wan 2.7 series is a comprehensive open-weight AI suite for image generation/editing and video creation, featuring thinking mode reasoning, first/last frame control, up to 4K images and 1080p videos, native audio sync, and exceptional text rendering accuracy.
8 モデル
Alibaba Cloud is a leading provider of advanced AI models, featuring the Qwen series (including Qwen-Image for multimodal vision-language tasks) and the Wan series for high-fidelity video generation.
65 モデル
ByteDance is a leading provider of advanced AI media models, featuring the Seedance series for high-fidelity multimodal video generation and the Seedream series for superior image creation and editing.
24 モデル
Google is a leading provider of advanced AI media models, featuring Nano Banana and Imagen for high-fidelity image generation and editing, and Veo for scalable video synthesis.
28 モデル
Kuaishou is a leading provider of advanced AI media models, featuring the Kling series (including Video and Image) for high-fidelity multimodal video and image generation.
19 モデル
Microsoft is a global technology company founded by Bill Gates and Paul Allen in 1975, best known for its Windows operating system, Office productivity suite, Azure cloud platform, and its mission to "empower every person and every organization on the planet to achieve more."
5 モデル
MiniMax is a leading provider of advanced AI media models, featuring Hailuo for high-fidelity multimodal video generation and editing.
18 モデル
OpenAI is an AI research and deployment company founded in 2015, dedicated to developing safe and beneficial artificial general intelligence (AGI) that benefits all of humanity.
5 モデル
PixVerse is an AI-powered video generation platform that transforms text prompts and images into high-quality, creative short videos with features like real-time interaction, AI effects, and one-click storytelling.
13 モデル
SkyWork is a versatile AI image generator that transforms text prompts into photorealistic, professional-grade visuals with lightning-fast speed, offering rich customization across artistic styles, colors, lighting, and composition in a seamless all-in-one creative workspace.
10 モデル
Vidu is a cutting-edge AI video generation model that transforms text prompts, images, and reference videos into high-quality cinematic content. It supports multiple modes including text-to-video and image-to-video, delivering fast generation with strong motion consistency and 1080p resolution.
23 モデル
xAI, founded by Elon Musk in 2023, is an artificial intelligence company best known for its Grok chatbot — a witty, rebellious AI integrated into the X platform — with a stated mission to "understand the true nature of the universe."
12 モデルModellixでテキストから画像モデルとAPIを検索できます。
34 モデルModellixで画像編集モデルとAPIを検索できます。
35 モデルModellixでテキストから動画モデルとAPIを検索できます。
35 モデルModellixで画像から動画モデルとAPIを検索できます。
73 モデルModellixで動画から動画モデルとAPIを検索できます。
29 モデルModellixでテキスト読み上げモデルとAPIを検索できます。
9 モデルModellixで音声認識モデルとAPIを検索できます。
5 モデルModellixで音声変換モデルとAPIを検索できます。
2 モデル
The best animation generation Models.
14 モデル
The best models for generating manga/anime images or videos.
14 モデル
AI models suitable for art design.
9 モデル
The best avatar-generating AI models.
9 モデル
The best portrait-generating AI models.
8 モデル
This is a collection of the best models for style transfer, including image generation and video generation models.
14 モデル
It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.
25 モデル
The mainstream image-colorization AI models.
7 モデル
This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.
2 モデル
The mainstream face-swap models.
6 モデル
The mainstream image upscalers.
6 モデル
A curated collection of lip-sync AI models.
3 モデル
The best LOGO generation design and creation AI models.
8 モデル
The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.
4 モデル
An AI models for restoring old photos.
7 モデル
Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.
2 モデル
The best voice-cloning models.
2 モデル
透明性の高い料金、Playgroundでの試用、本番ワークロード向けの統合APIで、画像・動画・音声モデルを利用できます。
検索を開始