在 Modellix 探索 38 个生产可用的文生视频 AI 模型,对比模型能力、在线试用,并通过统一 API 快速完成集成。

[Core Function] Gemini Omni Flash T2V is Google's fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.

**[Core Function]** SkyReels Text-to-Video generates a video purely from a text prompt, with no input media. **[Strengths]** Strong prompt adherence and smooth motion; supports optional audio, 480p/720p/1080p output, and a fast/std quality-speed trade-off. **[Best For]** Turning an idea or script into video, concept visualization, story beats, and social clips generated from text. **[Limitations]** Do NOT use this when the user provides an image, video, or audio input; route to Image-to-Video (skyreels-i2v), Reference-to-Video (skyreels-r2v), or the video-to-video / Omni models instead. Output is capped at 1080p and 15s per clip; fast mode currently supports only sound=false (no audio). **[Routing]** Use for text-only generation. Choose mode=fast for quicker results or mode=std for balanced quality; set resolution and aspect_ratio as needed.

[Core Function] PixVerse c1 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio. [Best For] Turning an idea or script into video, concept visualization, story beats, social clips from text. [Limitations] Do NOT use this when the user provides a starting image, two frames, or reference images; route to Image-to-Video, First-Last-Frame, or Reference-to-Video instead. c1 does not support multi-clip. [Routing] Use for text-only generation. The c1 variant takes the same inputs as v6 except it does not support multi-clip generation; choose c1-t2v when the user requests the c1 model and multi-clip is not needed.

[Core Function] PixVerse v6 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio and multi-clip generation. [Best For] Turning an idea or script into video, concept visualization, story beats, social clips from text. [Limitations] Do NOT use this when the user provides a starting image, two frames, or reference images; route to Image-to-Video, First-Last-Frame, or Reference-to-Video instead. [Routing] Use for text-only generation. v6 additionally supports multi-clip generation (generate_multi_clip_switch), which c1 does not; choose v6-t2v when multi-clip output is needed or the v6 model is requested.

[Core Function] Seedance 2.0 Mini T2V is the lightweight text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use for cost-efficient text-to-video.

[Core Function] Seedance 2.0 Fast T2V is the faster text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use when speed is preferred over maximum resolution.

[Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.

[Core Function] HappyHorse 1.1 T2V is Alibaba's latest streamlined text-to-video model. [Strengths] It generates 720P/1080P video with native audio support, 3-15 second duration, and an expanded set of aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: fast HappyHorse text-to-video generation, social video formats, and high-throughput content creation. [Limitations] It does not expose custom audio controls; use Wan 2.7 T2V when custom audio input is required. [Routing] Prefer this model when the user explicitly requests HappyHorse text-to-video or wants the latest HappyHorse generation quality.

[Core Function] Grok Imagine Video is xAI's text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social clips, animated concepts, and dynamic scene generation from a description. [Limitations] Do NOT use this model when you have a starting image or reference subjects, or when you need resolutions above 720p or clips longer than 15 seconds; it is limited to 480p/720p and 15s. [Routing] Use this model when the user wants a video from text only. If a starting image is provided, route to the Image-to-Video model; for reference-driven character video, use Reference-to-Video.

[Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot 'smart cuts', complex physics, and simultaneous audio-visual generation. [Best For] Highly recommended for: cinematic storytelling, professional advertising, short films, and high-fidelity concept visualizations. [Limitations] Do NOT use this model if the user is looking for an instant, low-latency preview, as generation takes longer. It does not support automatic BGM addition. [Routing] Use this model by default for high-quality text-to-video requests. If the user specifically asks for 'fast' or 'quick' generation, switch to the Q3 Turbo T2V model.

[Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, generating background b-roll, and scenarios where generation speed is prioritized. [Limitations] Do NOT use this model if you need ultimate cinematic quality, complex audio-visual synchronization, or multi-shot 'smart cuts'. It does not support automatic BGM addition. [Routing] Choose this 'Turbo' model when the user emphasizes 'quick', 'fast', or needs immediate results. If the user demands the highest cinematic quality or advanced audio-visual features, choose the Q3 Pro T2V model instead.

[Core Function] Wan 2.7 T2V is Alibaba's flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: high-quality commercial video generation, professional storytelling, and dynamic cinematic sequences. [Limitations] Do NOT use this model if the user specifically requests the streamlined 'HappyHorse' workflow. [Routing] Use this model by default for high-end text-to-video requests on the Alibaba platform.

[Core Function] Veo 3.1 Lite T2V is a balanced text-to-video model. [Strengths] It provides a good balance between generation speed and visual quality, still supporting the advanced architecture of the 3.1 series. [Best For] Highly recommended for: general video content creation and social media posts where 4K is not strictly necessary. [Limitations] Do NOT use this model for the absolute highest cinematic fidelity (use the standard Veo 3.1 instead). [Routing] Route to this model for standard, everyday video generation tasks.

[Core Function] HappyHorse 1.0 T2V is a breakout, highly optimized text-to-video model. [Strengths] It provides streamlined, fast, and high-quality video generation (up to 15s at 1080p) with native audio support, acting as a highly efficient alternative to Wan 2.7. [Best For] Highly recommended for: fast experimentation, rapid content creation, and users specifically requesting 'HappyHorse'. [Limitations] Might lack some of the deeply integrated legacy editing features found strictly within the broader Wan 2.7 ecosystem. [Routing] Route to this model when the user explicitly mentions 'HappyHorse' or desires a streamlined, high-performance alternative to Wan.

[Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.

[Core Function] MiniMax T2V-01 is a legacy text-to-video model. [Strengths] Standard text-to-video generation maintained for backward compatibility. [Best For] Recommended only for: maintaining existing integrations that specifically require the T2V-01 endpoint. [Limitations] Do NOT use this model for new creations. It lacks the physical realism of Hailuo 02 and the human performance/stylization of Hailuo 2.3. [Routing] Only use this if the user explicitly requests the legacy 'T2V-01' model. Otherwise, default to Hailuo 2.3.

[Core Function] MiniMax T2V-01-Director is a legacy text-to-video model with explicit camera controls. [Strengths] It natively accepts explicit camera movement commands (pan, zoom, tilt) alongside the text prompt. [Best For] Recommended for: legacy workflows that specifically rely on the Director-mode camera parameters. [Limitations] Do NOT use this model for general video generation, as it is an older architecture superseded by Hailuo 2.3 and 02. [Routing] Only use this if the user explicitly requests the 'Director' model or legacy 'T2V-01' generation. Otherwise, default to Hailuo 2.3.

[Core Function] Hailuo 02 T2V is a text-to-video generation model optimized for physical realism. [Strengths] It excels at complex physics simulation, fluid dynamics, broad cinematic scenes, and natively rendering 1080p video up to 10 seconds without downscaling. [Best For] Highly recommended for: product commercials, high-speed sports action, nature documentaries, and sweeping landscapes. [Limitations] Do NOT use this model for highly stylized anime/art or nuanced human micro-expressions, where Hailuo 2.3 performs better. [Routing] Route to this model when the user requests '1080p for 10 seconds', complex physical action (like splashing water or crashes), or broad landscapes. For human characters and stylization, use Hailuo 2.3 T2V.

[Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.

[Core Function] Veo 2 T2V is an older generation text-to-video model. [Strengths] Known for its distinctive cinematic style and fluid motion priors from the Veo 2 era. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Veo 3.1 T2V.

[Core Function] Veo 3 Fast T2V is an older generation text-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Veo 3.1 T2V.

[Core Function] Veo 3 T2V is an older generation text-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Veo 3.1 T2V.

[Core Function] Veo 3.1 Fast T2V is a high-speed text-to-video model. [Strengths] It is heavily optimized for fast generation, delivering video content with natively synchronized audio at 1080p quickly. [Best For] Highly recommended for: rapid prototyping, quick visual iteration, and high-volume background video generation. [Limitations] Do NOT use this model if you require 4K resolution or maximum artistic detail. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or 'rapid' video generation.

[Core Function] Veo 3.1 T2V is Google's state-of-the-art cinematic text-to-video engine. [Strengths] It natively generates 4K professional-grade video output with natively synchronized audio and supports complex camera movements. [Best For] Highly recommended for: high-end creative storytelling, cinematic short films, and experimental video production with sound. [Limitations] Do NOT use this model if you need instant/real-time generation, as 4K video rendering takes time. [Routing] Use this model by default for all high-quality text-to-video requests on the Google platform.

[Core Function] Kling V2.6 T2V is a stable, classic text-to-video model. [Strengths] It provides excellent semantic adherence, top stability, and supports 1080p generation with optional audio. [Best For] Highly recommended for: stable production workflows, 5s to 10s video generation, and scenarios requiring strict prompt adherence. [Limitations] Do NOT use this model if you need 4K resolution or durations longer than 10 seconds. It is superseded by V3 for high-end tasks. [Routing] Route to this model when users prefer the classic V2.6 aesthetic or specifically ask for a stable 1080p generation. Otherwise, use V3.

[Core Function] Kling V2.5 Turbo T2V is a fast, cost-effective text-to-video model. [Strengths] It offers very fast generation speeds and great value while maintaining 1080p resolution for 5s/10s clips. [Best For] Highly recommended for: rapid prototyping, batch social media creation, and cost-sensitive pipelines. [Limitations] Do NOT use this model for the absolute highest cinematic fidelity or native audio synchronization. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or 'cost-effective' generation.

[Core Function] Kling V2.1 Master T2V is an older generation text-to-video model. [Strengths] Offered improved multi-subject tracking and better texture details over V1. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 T2V.

[Core Function] Kling V2 Master T2V is an older generation text-to-video model. [Strengths] Offered improved multi-subject tracking and better texture details over V1. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 T2V.

[Core Function] Kling V1.6 T2V is an older generation text-to-video model. [Strengths] Pioneered early realistic physics simulation in video generation. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 T2V.

[Core Function] Kling V1 T2V is an older generation text-to-video model. [Strengths] Pioneered early realistic physics simulation in video generation. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kling V3 T2V.

[Core Function] Seedance 1.5 Pro T2V is a joint audio-video text-to-video generation model. [Strengths] It accurately follows complex text instructions to generate high-quality video with synchronized audio. [Best For] Highly recommended for: standard text-to-video generation where strict prompt adherence and audio are required. [Limitations] Do NOT use this model if you need the advanced multi-shot or multimodal reference capabilities of the 2.0 architecture. [Routing] Use this model by default for ByteDance text-to-video tasks.

[Core Function] Seedance 1.0 Pro T2V is an older generation text-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Seedance 1.5 Pro T2V.

[Core Function] Seedance 1.0 Pro Fast T2V is an older generation text-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Seedance 1.5 Pro T2V.

[Core Function] Wanx 2.1 T2V Plus is an older generation text-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 T2V.

[Core Function] Wanx 2.1 T2V Turbo is an older generation text-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 T2V.

[Core Function] Wan 2.2 T2V Plus is an older generation text-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 T2V.

[Core Function] Wan 2.5 T2V Preview is an older generation text-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 T2V.

[Core Function] Wan 2.6 T2V is an older generation text-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 T2V.