Modellix Logo Aurora Nasdaq Full Gray Font SVG
モデル
料金CLI
ドキュメント
ブログ
Modellixについて
日本語
始める
モデル/検索
カテゴリー

すべてのモデルを探す

(222件)
bytedance/seedance-2.5-v2v
bytedance/seedance-2.5-v2v
video-to-video

[Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 4-30 second output and mp4 or mov containers. [Best For] Highly recommended for: editing existing clips, extending motion from a base video, multi-reference restyling, and prompt-driven composition that cites Video n or Image n. [Limitations] Do NOT use this model for text-only or first-frame-only workflows. video_urls is required. Do NOT send first_frame_image or last_frame_image. Do NOT use this model when the user requires 1080p or 4k output. Do NOT use this model for audio-only input. [Routing] Prefer this model for Seedance 2.5 edit, extension, and video-reference jobs. Use Seedance 2.5 T2V for text-only and Seedance 2.5 I2V for image-first generation.

$0.3450~$0.6900/sec
bytedance/seedance-2.5-i2v
bytedance/seedance-2.5-i2v
image-to-video

[Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 4-30 second duration, and mp4 or mov output. [Best For] Highly recommended for: animating a keyframe, first-to-last transitions, multi-image character consistency, and image-led storytelling on Seedance 2.5. [Limitations] Do NOT mix first_frame_image or last_frame_image with reference_images. Do NOT send video_urls on I2V. Do NOT use audio_urls alone. Do NOT use this model when the user requires 1080p or 4k output. [Routing] Prefer this model for Seedance 2.5 image-driven generation. Use Seedance 2.5 T2V for text-only requests and Seedance 2.5 V2V when reference video is required.

$0.1184~$0.2656/sec
bytedance/seedance-2.5-t2v
bytedance/seedance-2.5-t2v
text-to-video

[Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-to-video storytelling, social clips beyond 15 seconds, and Seedance workflows that need mov output. [Limitations] Do NOT use this model if the user provides images, video, or audio as inputs. Do NOT use this model when the user requires 1080p or 4k output. [Routing] Prefer Seedance 2.5 T2V when the user needs more than 15 seconds of text-to-video. Use Seedance 2.0 T2V when 1080p or 4k is required. Use Seedance 2.5 I2V or V2V when media inputs are provided.

$0.1184~$0.2656/sec
alibaba/qwen-image-3.0-edit
alibaba/qwen-image-3.0-edit
image-to-image

[Core Function] Qwen Image 3.0 Edit is the standard image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite (direct mode), and 1-6 outputs while preserving subject identity. [Best For] Background replacement, outfit or style changes, multi-image fusion, and iterative retouching when Pro-tier quality is not required. [Limitations] Do NOT use this for pure text-to-image with no reference images (use Qwen Image 3.0 instead). Keep output total pixels within 512*512 to 2048*2048. prompt_extend_mode only supports direct (agent is T2I-only). [Routing] Route here when the user provides reference image(s) and wants balanced Qwen 3.0 edit quality. Prefer Qwen Image 3.0 Pro Edit for higher quality edits.

$0.0040/img
alibaba/qwen-image-3.0
alibaba/qwen-image-3.0
text-to-image

[Core Function] Qwen Image 3.0 is Alibaba's standard text-to-image model balancing quality and speed. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite (direct/agent modes), batch generation of 1-6 images, and long structured prompts. [Best For] General creative stills, posters with readable text, product shots, and multi-variant exploration (n up to 6) when Pro-tier photorealism is not required. [Limitations] Do NOT use this if the user needs image editing with reference images (use Qwen Image 3.0 Edit). Keep total pixels within 512*512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer Qwen Image 3.0 Pro for higher photorealism; use this for balanced quality/speed. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Edit instead.

$0.0020/img
kling/kling-v3-turbo-i2v
kling/kling-v3-turbo-i2v
image-to-video

[Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.

$0.0773~$0.0966/sec
kling/kling-v3-turbo-t2v
kling/kling-v3-turbo-t2v
text-to-video

[Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical 720p/1080p short videos rather than maximum cinematic headroom. [Best For] Highly recommended for: rapid prototyping, social and ad iteration, batch short-form pipelines, and dialogue clips where turnaround time and unit cost matter most. [Limitations] Do NOT use this model if the user requires peak 4K cinematic fidelity, heavy multi-shot storyboard control, or maximum visual polish; use Kling V3 T2V or Kling V3 Omni T2V instead. Do NOT use it for image-conditioned animation; use Kling V3 Turbo I2V or Kling V3 I2V. [Routing] Choose Kling V3 Turbo T2V when the user says fast, quick, cheap, or high volume. Otherwise default to Kling V3 T2V for quality, or Kling V3 Omni T2V when consistency and Omni-class control are requested.

$0.0773~$0.0966/sec
kling/kling-v3-omni-i2v
kling/kling-v3-omni-i2v
image-to-video

[Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.

$0.0580~$0.2898/sec
kling/kling-v3-omni-t2v
kling/kling-v3-omni-t2v
text-to-video

[Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and better adherence when scenes demand coherent characters or multi-beat storytelling from text alone. [Best For] Highly recommended for: narrative T2V with recurring subjects, dialogue-aware scenes, brand or product continuity across beats, and premium short films where consistency matters more than raw throughput. [Limitations] Do NOT use this model if the user only needs the cheapest or fastest clip; prefer Kling V3 Turbo T2V. Do NOT use it when the workflow is image-first or needs multi-image references; use Kling V3 Omni I2V or Kling V3 I2V instead. Do NOT use it for deep physics-reasoning specialty tasks better served by Kling Video O1. [Routing] Choose Kling V3 Omni T2V when the user emphasizes Omni, consistency, multimodal quality, or complex text narratives. Prefer Kling V3 T2V as the default high-quality T2V baseline; prefer Kling V3 Turbo T2V when the user stresses speed, cost, or high-volume short-form output.

$0.0580~$0.2898/sec
...
必要なモデルが見つかりませんか? ご要望をお聞かせください。
Modellix Logo Full Colored

あらゆるAIメディア生成を、1つの統合APIで。

プロダクト
モデルモデルを探す注目モデル
リソース
ドキュメントブログ料金CLI
会社情報
Modellixについてサービス利用規約プライバシーポリシー
シリーズ
CosyVoiceFunGemini OmniGPT ImageGrok ImagineGrok VoiceHailuo 02Hailuo 2.3HappyHorseImagenKling V3MAI ImageMiniMax Speech
Nano BananaPixVerse C1PixVerse V6Qwen AudioQwen ImageSeedanceSeedreamSkyReelsVeo 3Veo 3.1Vidu Q3Wan 2.6Wan 2.7
プロバイダー
AlibabaBytedanceGoogleKlingMicrosoftMiniMaxOpenAIPixVerseReveSkyworkViduxAI
コレクション
AI Animation GeneratorAI Anime GeneratorAI Art GeneratorAI Avatar GeneratorAI Portrait GeneratorAI Style TransferAI Video GeneratorColorize PhotoDigital HumanFace SwapImage UpscaleLip SyncLogo Generator
One ClickPhoto RestorationVirtual Try-OnVoice Cloning
カテゴリー
テキストから画像画像編集テキストから動画画像から動画動画から動画テキスト読み上げ音声認識音声変換

Copyright © 2026 METAVERSE CLOUD PTE. LTD. All rights reserved.

所在地: 60 Paya Lebar Road #12-03, Paya Lebar Square, Singapore 409051

メール: support@modellix.ai

  • English
  • 简体中文
  • 日本語
最新リリース
最新リリースはまだありません
画像 · モデル68件
GPT Image 2NEWNano Banana 2HOTSeedream 5.0 LiteHOT
画像モデルをすべて表示
動画 · モデル138件
Wan2.7 T2VNEWVeo 3.1 Lite T2VNEWHailuo 2.3 T2VHOTSeedance 2.0 I2VHOTHappyhorse 1.0 T2VHOT
動画モデルをすべて表示
注目
Seedance 2.5 T2VSeedream 5.0 ProGemini Omni Flash T2VV6 T2VHappyhorse 1.1 T2V
220件以上のモデルをすべて見る
はじめに
AI OnboardingQuick StartModel ProvidersPricing
利用方法
REST APICLIMCPAgent Skill
リソース
Product UpdatesModel UpdatesSupportDiscord Community
ドキュメントをすべて見る
トレンド
GPT Image 2 API Guide: Pricing & VariantsSeedream API Guide: 4K ImagesKling AI API: Pricing & IntegrationSeedance 2.0 API: Global Access
最新記事
HappyHorse-1.0: Open-Source AI Video ModelGoogle AI Suite Lands on ModellixModellix CLI: Generate from TerminalIntroducing Modellix
すべての記事を見る