
[Core Function] Vidu I2V New NS is an Image-to-Video model that animates a single first-frame image into an audio-visual synchronized, multi-shot narrative clip via the Vidu template API (vidu_i2v_new_ns). [Strengths] It excels at newer visual quality with multi-shot storytelling and longer durations from 2 to 15 seconds at 720p or 1080p (default 1080p). [Best For] Highly recommended for: cinematic storyboard-style clips from one still, brand films needing multi-shot continuity, longer narrative shorts up to 15s, and high-resolution audio-synced animation. [Limitations] Do NOT use this model if you need 360p or durations outside 2-15 seconds; it accepts only one input image and does not expose seed or aspect_ratio. [Routing] Choose Vidu I2V New NS for the latest multi-shot narrative I2V quality and longer clips. If the user needs 360p previews or only 3-10s with the q2-ns-duration template, route to Vidu Q2-NS instead.

[Core Function] Vidu Q2-NS is an Image-to-Video model that animates a single first-frame image into an audio-visual synchronized clip via the Vidu template API (q2-ns-duration). [Strengths] It excels at producing sound-synced motion from one still image with flexible durations from 3 to 10 seconds and resolutions including 360p for lower-cost previews. [Best For] Highly recommended for: short social clips with ambient audio, product motion from a hero still, cost-sensitive audio-visual drafts at 360p, and general first-frame animation with negative-prompt control. [Limitations] Do NOT use this model for multi-shot narrative storytelling or when you need durations outside 3-10 seconds; it accepts only one input image and does not expose seed or aspect_ratio. [Routing] Choose Vidu Q2-NS for audio-synced I2V with 360p option and 3-10s duration. If the user needs longer clips up to 15s, multi-shot narrative, or defaults to 1080p cinematic quality, route to Vidu I2V New NS instead.

[Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.

[Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical 720p/1080p short videos rather than maximum cinematic headroom. [Best For] Highly recommended for: rapid prototyping, social and ad iteration, batch short-form pipelines, and dialogue clips where turnaround time and unit cost matter most. [Limitations] Do NOT use this model if the user requires peak 4K cinematic fidelity, heavy multi-shot storyboard control, or maximum visual polish; use Kling V3 T2V or Kling V3 Omni T2V instead. Do NOT use it for image-conditioned animation; use Kling V3 Turbo I2V or Kling V3 I2V. [Routing] Choose Kling V3 Turbo T2V when the user says fast, quick, cheap, or high volume. Otherwise default to Kling V3 T2V for quality, or Kling V3 Omni T2V when consistency and Omni-class control are requested.

[Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.

[Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and better adherence when scenes demand coherent characters or multi-beat storytelling from text alone. [Best For] Highly recommended for: narrative T2V with recurring subjects, dialogue-aware scenes, brand or product continuity across beats, and premium short films where consistency matters more than raw throughput. [Limitations] Do NOT use this model if the user only needs the cheapest or fastest clip; prefer Kling V3 Turbo T2V. Do NOT use it when the workflow is image-first or needs multi-image references; use Kling V3 Omni I2V or Kling V3 I2V instead. Do NOT use it for deep physics-reasoning specialty tasks better served by Kling Video O1. [Routing] Choose Kling V3 Omni T2V when the user emphasizes Omni, consistency, multimodal quality, or complex text narratives. Prefer Kling V3 T2V as the default high-quality T2V baseline; prefer Kling V3 Turbo T2V when the user stresses speed, cost, or high-volume short-form output.

[Core Function] Kling Video O1 I2V is the image-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced generation from 1-7 reference images, 720p/1080p, duration 3-10s (single image only 5 or 10). [Best For] Complex physical motion grounded in reference frames. [Limitations] No native audio; no aspect_ratio (follows first frame); duration capped at 10s. [Routing] Prefer this for O1-quality I2V; use Kling Video O1 V2V when a source video is required.

[Core Function] Kling Video O1 T2V is the text-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced prompt planning with 3-10s duration and 720p/1080p output. [Best For] Complex physical interactions and logically demanding scenes from text alone. [Limitations] No native audio; duration capped at 10s; no multi_shot. [Routing] Prefer this for O1-quality T2V; use Kling Video O1 V2V when a source video is required.

[Core Function] Qwen-Audio 3.0 TTS Flash is Alibaba's low-latency Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Fast synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanhuan_v3.6, longjielidou_v3.6, loongeva_v3.6, and loongjohn (see the Qwen-Audio-TTS voice list). [Best For] Voice assistants, interactive prompts, short announcements, multilingual product flows, and latency-sensitive batch TTS using Qwen-Audio voices. [Limitations] Do NOT mix Plus-only voices (e.g. longanlingxin) with Flash. Use Plus when maximum narration quality matters more than turnaround time. [Routing] Choose Flash when speed matters most. Choose qwen-audio-3.0-tts-plus for premium narration quality.