Last updated
Alibaba's latest video generation model.
Generates video and synchronized audio together in a single pass with native multilingual lip-sync.
Available in both generation mode (text-to-video, image-to-video, reference-to-video) and video edit mode.
Best suited for dialogue-driven and performance scenes where audio-visual sync matters. Prefer Seedance 2.0 or Kling 3.0 for stronger realism, or 4K delivery.
Native audio in one pass: Generates dialogue, ambience, music, and Foley alongside the video itself, so sound stays in sync with motion without a separate audio pipeline.
Multilingual lip-sync: Matches mouth shapes to spoken phonetics across seven languages (English, Mandarin, Cantonese, Japanese, Korean, German, French).
Reference images: Takes up to nine reference images and keeps consistent faces, wardrobe, and identity across a cast, suited to multi-character and ensemble scenes.
Generation + edit modes: Works as a generation model (text/image/reference-to-video) and as a video editor, for refining existing clips rather than starting over.
Nine aspect ratios: Covers cinematic 21:9 through vertical 9:16 and square, plus 9:21, 5:4, and 4:5.
Resolution cap: Tops out at 1080p, no 4K tier, unlike Seedance 2.0 or Kling 3.0.
Last updated

