
[Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-to-video storytelling, social clips beyond 15 seconds, and Seedance workflows that need mov output. [Limitations] Do NOT use this model if the user provides images, video, or audio as inputs. Do NOT use this model when the user requires 1080p or 4k output. [Routing] Prefer Seedance 2.5 T2V when the user needs more than 15 seconds of text-to-video. Use Seedance 2.0 T2V when 1080p or 4k is required. Use Seedance 2.5 I2V or V2V when media inputs are provided.

[Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.

[Core Function] Qwen Image 3.0 Pro is Alibaba's latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and long structured prompts for complex layouts. [Best For] Highly recommended for: photorealistic stills, marketing posters with readable text, detailed scene compositions, multi-panel layouts, product hero shots, and multi-variant creative exploration (n up to 6). [Limitations] Do NOT use this if the user needs native 4K output, thinking-mode reasoning, or image editing with reference images (use Qwen Image 3.0 Pro Edit for edits). Keep total pixels within 512*512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer this over Qwen Image 2.0 Pro for new Qwen Image text-to-image work. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Pro Edit instead.

[Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.

[Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.

[Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanlingxin and longanlufeng (see the Qwen-Audio-TTS voice list). [Best For] Premium narration, brand voiceovers, multilingual product audio, and quality-sensitive batch TTS when Qwen-Audio voices are preferred. [Limitations] Do NOT mix Flash-only voices (e.g. longanhuan_v3.6) with Plus. [Routing] Choose Plus when speech quality is the priority. Choose qwen-audio-3.0-tts-flash when lower latency matters more.

[Core Function] Grok Imagine Image (Quality) is xAI's high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed concept art, marketing visuals, and any scenario where image quality is prioritized over generation speed. [Limitations] Do NOT use this model when latency is critical, as generation is slower than the standard model. [Routing] Use this model by default when the user emphasizes quality or detail. For faster, lighter generation use Grok Imagine Image (standard).

[Core Function] Wan 2.7 T2V is Alibaba's flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: high-quality commercial video generation, professional storytelling, and dynamic cinematic sequences. [Limitations] Do NOT use this model if the user specifically requests the streamlined 'HappyHorse' workflow. [Routing] Use this model by default for high-end text-to-video requests on the Alibaba platform.

[Core Function] Grok Imagine Video is xAI's text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social clips, animated concepts, and dynamic scene generation from a description. [Limitations] Do NOT use this model when you have a starting image or reference subjects, or when you need resolutions above 720p or clips longer than 15 seconds; it is limited to 480p/720p and 15s. [Routing] Use this model when the user wants a video from text only. If a starting image is provided, route to the Image-to-Video model; for reference-driven character video, use Reference-to-Video.

[Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 4-30 second output and mp4 or mov containers. [Best For] Highly recommended for: editing existing clips, extending motion from a base video, multi-reference restyling, and prompt-driven composition that cites Video n or Image n. [Limitations] Do NOT use this model for text-only or first-frame-only workflows. video_urls is required. Do NOT send first_frame_image or last_frame_image. Do NOT use this model when the user requires 1080p or 4k output. Do NOT use this model for audio-only input. [Routing] Prefer this model for Seedance 2.5 edit, extension, and video-reference jobs. Use Seedance 2.5 T2V for text-only and Seedance 2.5 I2V for image-first generation.

[Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 4-30 second duration, and mp4 or mov output. [Best For] Highly recommended for: animating a keyframe, first-to-last transitions, multi-image character consistency, and image-led storytelling on Seedance 2.5. [Limitations] Do NOT mix first_frame_image or last_frame_image with reference_images. Do NOT send video_urls on I2V. Do NOT use audio_urls alone. Do NOT use this model when the user requires 1080p or 4k output. [Routing] Prefer this model for Seedance 2.5 image-driven generation. Use Seedance 2.5 T2V for text-only requests and Seedance 2.5 V2V when reference video is required.

[Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-to-video storytelling, social clips beyond 15 seconds, and Seedance workflows that need mov output. [Limitations] Do NOT use this model if the user provides images, video, or audio as inputs. Do NOT use this model when the user requires 1080p or 4k output. [Routing] Prefer Seedance 2.5 T2V when the user needs more than 15 seconds of text-to-video. Use Seedance 2.0 T2V when 1080p or 4k is required. Use Seedance 2.5 I2V or V2V when media inputs are provided.

[Core Function] Qwen Image 3.0 Edit is the standard image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite (direct mode), and 1-6 outputs while preserving subject identity. [Best For] Background replacement, outfit or style changes, multi-image fusion, and iterative retouching when Pro-tier quality is not required. [Limitations] Do NOT use this for pure text-to-image with no reference images (use Qwen Image 3.0 instead). Keep output total pixels within 512*512 to 2048*2048. prompt_extend_mode only supports direct (agent is T2I-only). [Routing] Route here when the user provides reference image(s) and wants balanced Qwen 3.0 edit quality. Prefer Qwen Image 3.0 Pro Edit for higher quality edits.

[Core Function] Qwen Image 3.0 is Alibaba's standard text-to-image model balancing quality and speed. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite (direct/agent modes), batch generation of 1-6 images, and long structured prompts. [Best For] General creative stills, posters with readable text, product shots, and multi-variant exploration (n up to 6) when Pro-tier photorealism is not required. [Limitations] Do NOT use this if the user needs image editing with reference images (use Qwen Image 3.0 Edit). Keep total pixels within 512*512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer Qwen Image 3.0 Pro for higher photorealism; use this for balanced quality/speed. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Edit instead.

[Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.

[Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical 720p/1080p short videos rather than maximum cinematic headroom. [Best For] Highly recommended for: rapid prototyping, social and ad iteration, batch short-form pipelines, and dialogue clips where turnaround time and unit cost matter most. [Limitations] Do NOT use this model if the user requires peak 4K cinematic fidelity, heavy multi-shot storyboard control, or maximum visual polish; use Kling V3 T2V or Kling V3 Omni T2V instead. Do NOT use it for image-conditioned animation; use Kling V3 Turbo I2V or Kling V3 I2V. [Routing] Choose Kling V3 Turbo T2V when the user says fast, quick, cheap, or high volume. Otherwise default to Kling V3 T2V for quality, or Kling V3 Omni T2V when consistency and Omni-class control are requested.

[Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.

CosyVoice is a family of open-source TTS models by FunAudioLLM that delivers high-quality speech synthesis, zero-shot voice cloning, and low-latency streaming from v1.0 to v3.0.
4 个模型
Fun is Alibaba's open-source, end-to-end automatic speech recognition toolkit supporting multilingual ASR, voice activity detection, punctuation restoration, and speaker diarization with real-time streaming capabilities.
2 个模型
Gemini Omni is Google's multimodal video generation and editing model that lets you create, remix, and edit videos as easily as having a conversation — blending text, images, and video input with natural language commands.
4 个模型
The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.
4 个模型
Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.
10 个模型
Grok Voice is xAI's native speech-to-speech model powering expressive, real-time audio interactions with sub-second latency and agentic tool capabilities.
2 个模型
MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.
3 个模型
MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.
3 个模型
HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.
7 个模型
Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.
10 个模型
Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.
4 个模型
Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.
8 个模型
PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.
4 个模型
PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.
5 个模型
Qwen-Audio is a unified audio-language model series by Alibaba Cloud that processes speech, natural sounds, music, and singing across multiple languages and tasks, enabling universal audio understanding and multimodal interaction.
2 个模型
Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.
16 个模型
ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.
18 个模型
ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.
9 个模型
SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.
4 个模型
Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.
4 个模型
Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.
6 个模型
Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.
12 个模型
Alibaba's Wan 2.6 is a powerful open-source AI video generation model that creates cinematic 1080p multi-shot videos with native audio-visual synchronization, supporting text-to-video, image-to-video, and professional storytelling workflows.
7 个模型
Alibaba's Wan 2.7 series is a comprehensive open-weight AI suite for image generation/editing and video creation, featuring thinking mode reasoning, first/last frame control, up to 4K images and 1080p videos, native audio sync, and exceptional text rendering accuracy.
8 个模型
Alibaba Cloud is a leading provider of advanced AI models, featuring the Qwen series (including Qwen-Image for multimodal vision-language tasks) and the Wan series for high-fidelity video generation.
67 个模型
ByteDance is a leading provider of advanced AI media models, featuring the Seedance series for high-fidelity multimodal video generation and the Seedream series for superior image creation and editing.
27 个模型
Google is a leading provider of advanced AI media models, featuring Nano Banana and Imagen for high-fidelity image generation and editing, and Veo for scalable video synthesis.
25 个模型
Kuaishou is a leading provider of advanced AI media models, featuring the Kling series (including Video and Image) for high-fidelity multimodal video and image generation.
19 个模型
Microsoft is a global technology company founded by Bill Gates and Paul Allen in 1975, best known for its Windows operating system, Office productivity suite, Azure cloud platform, and its mission to "empower every person and every organization on the planet to achieve more."
5 个模型
MiniMax is a leading provider of advanced AI media models, featuring Hailuo for high-fidelity multimodal video generation and editing.
18 个模型
OpenAI is an AI research and deployment company founded in 2015, dedicated to developing safe and beneficial artificial general intelligence (AGI) that benefits all of humanity.
5 个模型
PixVerse is an AI-powered video generation platform that transforms text prompts and images into high-quality, creative short videos with features like real-time interaction, AI effects, and one-click storytelling.
13 个模型
SkyWork is a versatile AI image generator that transforms text prompts into photorealistic, professional-grade visuals with lightning-fast speed, offering rich customization across artistic styles, colors, lighting, and composition in a seamless all-in-one creative workspace.
10 个模型
Vidu is a cutting-edge AI video generation model that transforms text prompts, images, and reference videos into high-quality cinematic content. It supports multiple modes including text-to-video and image-to-video, delivering fast generation with strong motion consistency and 1080p resolution.
21 个模型
xAI, founded by Elon Musk in 2023, is an artificial intelligence company best known for its Grok chatbot — a witty, rebellious AI integrated into the X platform — with a stated mission to "understand the true nature of the universe."
12 个模型
The best animation generation Models.
15 个模型
The best models for generating manga/anime images or videos.
14 个模型
AI models suitable for art design.
9 个模型
The best avatar-generating AI models.
9 个模型
The best portrait-generating AI models.
6 个模型
This is a collection of the best models for style transfer, including image generation and video generation models.
14 个模型
It brings together the world's best video generation models, including text-to-video, image-to-video, and video editing capabilities.
27 个模型
The mainstream image-colorization AI models.
7 个模型
This page aggregates high-quality AI digital human generation models, which can help you easily create digital human videos, such as lip-syncing.
2 个模型
The mainstream face-swap models.
6 个模型
The mainstream image upscalers.
6 个模型
A curated collection of lip-sync AI models.
3 个模型
The best LOGO generation design and creation AI models.
6 个模型
The AI models here let you generate ready-to-use images or videos with just one click, covering applications such as marketing, advertising, short dramas, and more.
4 个模型
An AI models for restoring old photos.
7 个模型
Mainstream AI virtual try-on models—upload your model and clothing images to see how they fit.
2 个模型
The best voice-cloning models.
2 个模型