alibaba/happyhorse-1.0-r2v

happyhorse-1.0-r2v
Docs
Schema

[Core Function] HappyHorse 1.0 R2V is a reference-to-video model. [Strengths] It excels at maintaining character consistency using up to 9 reference images while generating new video actions based on a prompt. [Best For] Highly recommended for: character-consistent storytelling and generating multiple scenes with the same subject. [Limitations] Do NOT use if you simply want to animate a single image exactly as it is (use HappyHorse I2V instead). [Routing] Use this when the user provides reference images to dictate character/subject appearance in a newly generated action.

$0.1280~$0.2290/sec
image-to-video

Input

Text prompt describing the scene and referencing characters. Required. Use character1, character2, character3, etc. to reference images in the reference_images array (first image = character1, second = character2, etc.). Supports any language input. Maximum length: 5000 non-Chinese characters or 2500 Chinese characters (automatically truncated if exceeded)
Array of reference image URLs. Supports 1-9 images. Images are referenced in the prompt using character1, character2, etc., following array order. Image requirements: Format: JPEG, JPG, PNG, WEBP; Resolution: Short edge >= 400 pixels (720P+ recommended); File size: <= 10MB per image. Supports HTTP/HTTPS URLs
Hint: Drag and drop files, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Video duration in seconds. Must be an integer between 3 and 15
5
Video aspect ratio. Determines the output video dimensions
16:9
Video resolution level. The model automatically scales to the nearest total pixels based on the selected resolution
1080P
Random seed for reproducibility. If not specified, the system generates a random seed. Note: Due to the probabilistic nature of model generation, even with the same seed, results may not be completely identical

Result

No results yet

Run the model to preview the output here.

Next:

README

Parameters

Name Description Type Required Enums
prompt Text prompt describing the scene and referencing characters. Required. Use character1, character2, character3, etc. to reference images in the reference_images array (first image = character1, second = character2, etc.). Supports any language input. Maximum length: 5000 non-Chinese characters or 2500 Chinese characters (automatically truncated if exceeded) string Yes -
reference_images Array of reference image URLs. Supports 1-9 images. Images are referenced in the prompt using character1, character2, etc., following array order. Image requirements: Format: JPEG, JPG, PNG, WEBP; Resolution: Short edge >= 400 pixels (720P+ recommended); File size: <= 10MB per image. Supports HTTP/HTTPS URLs string[] Yes -
duration Video duration in seconds. Must be an integer between 3 and 15 integer No 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
ratio Video aspect ratio. Determines the output video dimensions string No 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9
resolution Video resolution level. The model automatically scales to the nearest total pixels based on the selected resolution string No 720P, 1080P
seed Random seed for reproducibility. If not specified, the system generates a random seed. Note: Due to the probabilistic nature of model generation, even with the same seed, results may not be completely identical integer No -

Pricing

Unit: $/sec

Dimension Pricing
resolution: 720P 0.1280
resolution: 1080P 0.2290
  • alibaba/happyhorse-1.0-t2v: [Core Function] HappyHorse 1.0 T2V is a breakout, highly optimized text-to-video model. [Strengths] It provides streamlined, fast, and high-quality video generation (up to 15s at 1080p) with native audio support, acting as a highly efficient alternative to Wan 2.7. [Best For] Highly recommended for: fast experimentation, rapid content creation, and users specifically requesting ‘HappyHorse’. [Limitations] Might lack some of the deeply integrated legacy editing features found strictly within the broader Wan 2.7 ecosystem. [Routing] Route to this model when the user explicitly mentions ‘HappyHorse’ or desires a streamlined, high-performance alternative to Wan.
  • alibaba/happyhorse-1.0-i2v: [Core Function] HappyHorse 1.0 I2V is a streamlined image-to-video model. [Strengths] It generates high-quality 720P/1080P video (3-15s) from an image efficiently, with native audio support. [Best For] Highly recommended for: rapid image animation and robust character motion. [Limitations] Does not support complex video continuation like Wan 2.7 I2V. [Routing] Use when the user requests ‘HappyHorse’ or a streamlined image animation.
  • alibaba/happyhorse-1.0-video-edit: [Core Function] HappyHorse 1.0 Video Edit is a streamlined video editing model. [Strengths] It provides high-quality video editing capabilities (with or without reference images) within the highly optimized HappyHorse architecture. [Best For] Highly recommended for: fast, high-quality video modifications, especially when the user explicitly requests HappyHorse. [Limitations] May lack the deep instruction-based logical replacement mechanics of Wan 2.7 Video Editing. [Routing] Route to this model when the user explicitly requests ‘HappyHorse’ for their video editing task.
  • alibaba/happyhorse-1.1-t2v: [Core Function] HappyHorse 1.1 T2V is Alibaba’s latest streamlined text-to-video model. [Strengths] It generates 720P/1080P video with native audio support, 3-15 second duration, and an expanded set of aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: fast HappyHorse text-to-video generation, social video formats, and high-throughput content creation. [Limitations] It does not expose custom audio controls; use Wan 2.7 T2V when custom audio input is required. [Routing] Prefer this model when the user explicitly requests HappyHorse text-to-video or wants the latest HappyHorse generation quality.
  • alibaba/happyhorse-1.1-i2v: [Core Function] HappyHorse 1.1 I2V is Alibaba’s latest streamlined first-frame image-to-video model. [Strengths] It turns a single image into high-quality 720P/1080P video with native audio support and 3-15 second duration; output aspect ratio follows the first frame image. [Best For] Highly recommended for: rapid image animation, product motion previews, and simple character or scene animation. [Limitations] It does not accept an explicit ratio parameter; use T2V or R2V when you need a fixed generated aspect ratio. [Routing] Prefer this model when the user provides one image and requests HappyHorse image animation.
  • alibaba/happyhorse-1.1-r2v: [Core Function] HappyHorse 1.1 R2V is Alibaba’s latest reference-image-to-video model. [Strengths] It uses 1-9 reference images to preserve subject or character appearance while generating new video actions, supports 720P/1080P output, 3-15 second duration, and expanded aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: character-consistent storytelling, reference-based product shots, and multi-image subject composition. [Limitations] Do NOT use if the user simply wants to animate a single image exactly as provided; use HappyHorse I2V instead. [Routing] Use when the user provides one or more reference images and asks for a newly generated HappyHorse video.

More in this series