For the complete documentation index, see llms.txt. This page is also available as Markdown.

Happy Horse 1.1

Alibaba's latest video generation model.

Overview

Generates video and synchronized audio together in a single pass with native multilingual lip-sync.

Available in both generation mode (text-to-video, image-to-video, reference-to-video) and video edit mode.

Best suited for dialogue-driven and performance scenes where audio-visual sync matters. Prefer Seedance 2.0 or Kling 3.0 for stronger realism, or 4K delivery.

Key updates

  • Native audio in one pass: Generates dialogue, ambience, music, and Foley alongside the video itself, so sound stays in sync with motion without a separate audio pipeline.

  • Multilingual lip-sync: Matches mouth shapes to spoken phonetics across seven languages (English, Mandarin, Cantonese, Japanese, Korean, German, French).

  • Reference images: Takes up to nine reference images and keeps consistent faces, wardrobe, and identity across a cast, suited to multi-character and ensemble scenes.

  • Generation + edit modes: Works as a generation model (text/image/reference-to-video) and as a video editor, for refining existing clips rather than starting over.

  • Nine aspect ratios: Covers cinematic 21:9 through vertical 9:16 and square, plus 9:21, 5:4, and 4:5.

Weaknesses

  • Resolution cap: Tops out at 1080p, no 4K tier, unlike Seedance 2.0 or Kling 3.0.

Last updated