text-to-video
Veo 3.1 T2V

[Core Function] Veo 3.1 T2V is Google's state-of-the-art cinematic text-to-video engine. [Strengths] It natively generates 4K professional-grade video output with natively synchronized audio and supports complex camera movements. [Best For] Highly recommended for: high-end creative storytelling, cinematic short films, and experimental video production with sound. [Limitations] Do NOT use this model if you need instant/real-time generation, as 4K video rendering takes time. [Routing] Use this model by default for all high-quality text-to-video requests on the Google platform.

探索 AI 模型

google/veo-3.1-i2v
google/veo-3.1-i2v
image-to-video

[Core Function] Veo 3.1 I2V is Google's cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animating concept art, creating cinematic transitions between images, and high-end video production. [Limitations] When using first+last frame or reference-only modes, the `duration` parameter must strictly be 8. Negative prompts are not supported in reference-only mode. [Routing] Use this model by default for high-quality image-to-video tasks or when multiple reference images are provided.

bytedance/seedream-5.0-lite-edit
bytedance/seedream-5.0-lite-edit
image-to-image

[Core Function] Seedream 5.0 Lite Edit is a reasoning-enhanced, smart image editing model. [Strengths] It features superior cross-modal understanding and reasoning, allowing for highly accurate, interactive multi-turn image editing with real-time knowledge enhancement. [Best For] Highly recommended for: complex image editing tasks, structural modifications, and edits requiring deep semantic understanding. [Limitations] As a 'Lite' model, raw visual rendering might not match the 4.5 tier. [Routing] Use this model by default for complex, reasoning-based image editing tasks.

google/nano-banana-2-edit
google/nano-banana-2-edit
image-to-image

[Core Function] Nano Banana 2 Edit is a high-speed image editing model. [Strengths] It rapidly modifies existing images or extracts image frames from videos based on text prompts. [Best For] Highly recommended for: rapid style transfer, quick image modifications, and fast creative edits. [Limitations] Do NOT use this model for meticulous photorealistic retouching. [Routing] Use this model by default for fast, creative image editing tasks.

bytedance/seedance-2.0-i2v
bytedance/seedance-2.0-i2v
image-to-video

[Core Function] Seedance 2.0 I2V is ByteDance's flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.

bytedance/seedance-2.0-v2v
bytedance/seedance-2.0-v2v
video-to-video

[Core Function] Seedance 2.0 V2V is ByteDance's flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] Highly recommended for: complex video-to-video transformations, restyling existing footage, and creating dynamic multi-shot edits. [Limitations] Do NOT use this model for simple still-image generation (use Seedream instead). [Routing] Use this model by default for any video editing or video-to-video generation tasks.

xai/grok-imagine-video-1.5-i2v
xai/grok-imagine-video-1.5-i2v
image-to-video

[Core Function] Grok Imagine Video 1.5 I2V animates a single starting image into a video using the Grok Imagine 1.5 generation backbone. [Strengths] It excels at producing motion from one starting frame with the improved 1.5 model. [Best For] Highly recommended for: animating a photo or illustration when the 1.5 generation backbone is preferred. [Limitations] Do NOT use this model for text-only generation, for resolutions above 1080p, or for clips longer than 15 seconds; it requires a starting image and is limited to 480p/720p/1080p and 15s. [Routing] Use this model when the user provides one starting image and prefers the 1.5 backbone.

xai/grok-imagine-image-quality-edit
xai/grok-imagine-image-quality-edit
image-to-image

[Core Function] Grok Imagine Image Edit (Quality) is xAI's high-fidelity image editing model. [Strengths] It excels at applying detailed, prompt-guided edits and style transformations to one or more source images while preserving fine detail. [Best For] Highly recommended for: high-quality restyling, detailed inpainting-style edits, and combining up to 3 source images. [Limitations] Do NOT use this model when latency is critical, as it is slower than the standard edit variant; a maximum of 3 source images is supported. [Routing] Use this model by default for quality-sensitive edits. For faster edits, route to Grok Imagine Image Edit (standard).

alibaba/wan2.7-videoedit
alibaba/wan2.7-videoedit
video-to-video

[Core Function] Wan 2.7 Video Editing is an instruction-based video modification model. [Strengths] It supports complex video editing tasks like content replacement using reference images, and replicating actions, effects, and camera movements. [Best For] Highly recommended for: modifying existing video footage, style transfer on videos, and targeted element replacement. [Limitations] Do NOT use this model to generate a brand new video from scratch; it requires an input video. [Routing] Use this model by default whenever a user wants to edit, alter, or restyle an existing video.

最新发布

skywork/sky-lipsync
skywork/sky-lipsync
video-to-video

**[Core Function]** SkyReels Lip Sync (retalking) re-drives a talking video so the subject's lips match a given audio track. **[Strengths]** Accurate lip re-synchronization on an existing talking-head video. **[Best For]** Dubbing, re-voicing talking-head video, and localizing spoken video. **[Limitations]** Do NOT use this to generate new motion or content from scratch; it only re-times lips on an existing video. Requires video_url and audio_url. Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and the target audio_url; optionally provide reference_char_url to guide the driven face.

skywork/video-extension-shot-switching
skywork/video-extension-shot-switching
video-to-video

**[Core Function]** SkyReels Shot-Switching Extension continues a video while transitioning to a new shot or camera angle. **[Strengths]** Cinematic shot transitions (cut-in, cut-out, reverse-shot, multi-angle, cut-away) when extending footage. **[Best For]** Adding a new shot after existing footage and cinematic transitions. **[Limitations]** Do NOT use this for a plain single-shot continuation (use video-extension-single-shot) or for generation from scratch. Requires a prefix_video (mp4 URL); adds 2-5s. **[Routing]** Provide prompt and prefix_video; choose cut_type for the transition style, or Auto to let the model decide.

skywork/video-extension-single-shot
skywork/video-extension-single-shot
video-to-video

**[Core Function]** SkyReels Single-Shot Extension continues an existing single-shot video, generating additional seconds guided by a text prompt. **[Strengths]** Seamless single-shot continuation of the existing motion and scene. **[Best For]** Lengthening clips and continuing an action within one continuous shot. **[Limitations]** Do NOT use this to create a video from scratch (use skyreels-t2v) or to switch shots / add transitions (use video-extension-shot-switching). Requires a prefix_video (mp4 URL); adds 5-30s. **[Routing]** Provide prompt and prefix_video; set duration for how many seconds (5-30) to append.

skywork/video-restyling
skywork/video-restyling
video-to-video

**[Core Function]** SkyReels Restyle re-renders an existing video into a preset visual style. **[Strengths]** Consistent style transfer across all frames into a chosen named art style. **[Best For]** Turning footage into simpsons, lego, paper-cutting, amigurumi, animal-crossing, van-gogh, or pixel-art looks. **[Limitations]** Do NOT use this to change content, motion, or add new scenes; it only restyles an existing video (input <=30s). Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and a style_name from the supported list.

skywork/skyreels-omni
skywork/skyreels-omni
video-to-video

**[Core Function]** SkyReels Omni is a reference-driven video model that generates or edits video using reference images (@image) and/or a reference video (@video), bound by tags in the prompt. **[Strengths]** A single endpoint covers motion reference, subject/background replacement, object insertion/removal, local editing, and video extension. **[Best For]** Video subject or background swap, motion transfer onto an image, object add/remove, local video edits, and extending a reference video. **[Limitations]** Do NOT use this for pure text-to-video (use skyreels-t2v) or simple single-image animation (use skyreels-i2v). Each ref tag must appear in the prompt as @tag; ref_videos supports only one video (<=15s); a reference-type video may combine only with image-type ref_images, while an extend-type video cannot combine with ref_images. When ref_videos is provided, aspect_ratio is ignored (output matches the video). **[Routing]** Provide ref_images (type grid or image) for image references and/or a single ref_videos entry (type reference for motion/edit, type extend for continuation); the reference tags must be used in the prompt.

skywork/segmented-camera-motion
skywork/segmented-camera-motion
image-to-video

**[Core Function]** SkyReels Segmented Camera Motion (audio-to-video) generates a talking-avatar video with directed camera movement across time segments. **[Strengths]** Combines an audio-driven avatar with per-segment camera trajectories such as push, pan, crane and rotation. **[Best For]** Dynamic presenter clips and cinematic avatar shots with controlled camera motion. **[Limitations]** Do NOT use this when you need a completely static camera (use single-actor-avatar). It requires first_frame_image and one audio segment; use camera_control_pro for multi-segment or compound moves. mode=std outputs 720p, mode=pro outputs 1080p. **[Routing]** Set a single traj_type plus camera_control_strength for a simple move, or supply camera_control_pro (a list of per-segment {start_time, end_time, traj_type, ...}) for compound motion; choose mode=pro for 1080p.

skywork/single-actor-avatar
skywork/single-actor-avatar
image-to-video

**[Core Function]** SkyReels Single-Actor Avatar (audio-to-video) drives a talking-avatar video from a single portrait image and one audio track. **[Strengths]** Lip-synced single-speaker talking-head video generated from an image plus audio. **[Best For]** Virtual presenters, single-speaker dubbing, and talking avatars. **[Limitations]** Do NOT use this for multi-speaker scenes (use the multi-actor flow) or when you only have text. It requires first_frame_image and exactly one audio segment (<=200s). mode=std outputs 720p, mode=pro outputs 1080p. **[Routing]** Provide a portrait first_frame_image and one audio URL in audios; choose mode=pro for 1080p output.

skywork/skyreels-r2v
skywork/skyreels-r2v
image-to-video

**[Core Function]** SkyReels Reference-to-Video (multiobject) generates a video from a prompt while preserving the subjects from 1-4 reference images. **[Strengths]** Multi-subject identity preservation, placing specific characters or objects into a newly generated scene. **[Best For]** Putting given characters/products into a new scene, multi-subject composition from reference photos. **[Limitations]** Do NOT use this to animate a single fixed frame (use skyreels-i2v) or for text-only generation (use skyreels-t2v). It requires 1-4 reference images and produces clips up to 5s. **[Routing]** Provide 1-4 subject reference images in ref_images; set aspect_ratio and duration (1-5s) as needed.

模型系列

GPT Image

The GPT-Image series by OpenAI consists of advanced multimodal models, such as GPT-Image-1 and GPT-Image-2, designed for generating and editing photorealistic images from text and image inputs.

Grok Imagine

Grok Imagine is xAI's cross-modal AI model series that unifies text-to-image, image-to-image, text-to-video, image-to-video, and video-to-video generation in a single visual system, delivering studio-grade, photorealistic visuals with best-in-class text rendering and precise creative control.

Hailuo 02

MiniMax's Hailuo 02 series is a top-ranked cinematic AI video suite for T2V/I2V, generating native 1080p clips with ultra-realistic physics, character consistency, and director-level controls.

Hailuo 2.3

MiniMax's Hailuo 2.3 series elevates cinematic AI video gen with 4K T2V/I2V, hyper-realistic physics/motion, extended clips, and advanced character consistency.

HappyHorse

HappyHorse is a leading open-source AI video generation model with 15 billion parameters that jointly produces high-quality 1080p videos and synchronized audio from text or image prompts, currently topping the Artificial Analysis Video Arena leaderboard.

Imagen

Google Imagen is Google's premier text-to-image diffusion model, excelling in photorealistic, high-resolution image generation from textual prompts with unmatched detail, creativity, and adherence to complex descriptions.

Kling V3

Kuaishou's Kling v3 series is an open multimodal AI suite for T2I/I2V/T2V, generating 4K cinematic visuals with native audio, multi-shot narratives, precise motion control, and consistent characters.

MAI Image

Microsoft's **MAI Image** series is a family of in-house, diffusion-based AI models, designed for state-of-the-art text-to-image generation and precise image-to-image editing, with a strong emphasis on photorealism, prompt adherence, and text rendering accuracy.

Nano Banana

Nano Banana is an advanced AI image generation and editing model based on Google's Gemini technology, delivering fast, precise transformations with exceptional prompt understanding, consistent character editing, and high-quality visuals.

PixVerse C1

PixVerse C1 is PixVerse's first AI video model purpose-built for film production, combining an industrial-grade action engine, cinematic VFX, storyboard-to-video conversion, and reference-guided character consistency to generate up to 15-second 1080p videos with native audio.

PixVerse V6

PixVerse V6 is PixVerse's flagship multi-shot AI video generation model that creates up to 15-second 1080p cinematic videos with native synchronized audio from text or image prompts, featuring improved camera control, consistent character emotion across scenes, and realistic physics simulation.

Qwen Image

Qwen Image is Alibaba's unified 7B text-to-image generation and editing model series, renowned for high-fidelity visuals, superior text rendering, Photoshop-like layered editing, and top rankings on global leaderboards.

Reve

Reve is a state-of-the-art text-to-image AI model known for its striking aesthetic quality, precise instruction following, and best-in-class typography rendering, consistently ranking among the top models on image generation leaderboards.

Seedance

ByteDance's Seedance is a multimodal AI video generation model that creates cinematic 1080p multi-shot videos from text, images, audio, or video prompts with immersive audio-visual realism and director-level creative controls.

Seedream

ByteDance's Seedream is a high-fidelity text-to-image and editing model supporting native 4K resolution, batch generation, superior typography, and consistent character rendering for professional creative workflows.

SkyReels

SkyReels is a powerful AI cinematic video generation model that transforms text and images into Hollywood-grade, human-centric videos with advanced facial animation, synchronized audio, and professional lighting — making it one of the leading open-source video foundation models available today.

Veo 3

Google Veo 3 is Google DeepMind's groundbreaking text-to-video AI model, unveiled at Google I/O 2025, that generates high-fidelity 4K cinematic videos with native synchronized audio from text or image prompts, offering professional controls and multi-scene coherence.

Veo 3.1

Google Veo 3.1 is the advanced successor to Veo 3, released in October 2025, enhancing 4K video generation with richer native audio, superior narrative control, precise image-to-video conversion, and seamless character consistency for dynamic storytelling.

Vidu Q3

Vidu Q3 is Shengshu AI’s advanced text-to-video and image-to-video model that generates up to 16-second clips with native audio, enhanced motion, and precise camera control.

Wan 2.6

Alibaba's Wan 2.6 is a powerful open-source AI video generation model that creates cinematic 1080p multi-shot videos with native audio-visual synchronization, supporting text-to-video, image-to-video, and professional storytelling workflows.

Wan 2.7

Alibaba's Wan 2.7 series is a comprehensive open-weight AI suite for image generation/editing and video creation, featuring thinking mode reasoning, first/last frame control, up to 4K images and 1080p videos, native audio sync, and exceptional text rendering accuracy.

接入领先 AI 媒体模型

探索图片、视频与音频模型,通过透明定价、在线试运行和统一 API 快速接入生产环境。

开始搜索
没有找到需要的模型?告诉我们。
视频生成140
图片生成79