For the complete documentation index, see llms.txt. This page is also available as Markdown.

Seedance 2.5

A state-of the art video generation model developed by ByteDance.

Overview

Evolution from Seedance 2.0.

Key upgrades include up to 30-second generations (double the previous ceiling), sharper realism, finer creative control with bold camera moves, massively expanded reference budgets, more languages, and negative prompts that actually hold.

The prompt is a script: write with timecodes, end states and transitions, and the model shoots them in one take.

Twice the price of Seedance 2.0.

Note: Content policies frequently block reference images containing humans, even innocuous ones. See Weaknesses below for workarounds.


Strengths

  • 30-second one-takes: Real narrative arcs in a single generation. Pair with video extension for up to 60s final output — minute-long ads in two passes.

  • Duration-adaptive scripts: Shortening the requested length compresses pacing instead of trimming beats. One master script declines cleanly into 30s / 20s / 12s / 8s cuts with no rewriting — the cheapest format-adaptation pipeline on any video model.

  • Massive reference budgets: Up to 30 image references, 10 videos (30s combined), and 10 audio clips (30s combined) per run.

  • Lip-synced dialogue in 11 languages: ZH, EN, ES, ID, MS most heavily optimized; TH, AR, PT, VI, JA, KO fully supported. On-screen text (signage, packaging, captions) can be forced to a target language.

  • High-precision editing: Replace, add, or remove elements in a time window while everything else stays untouched. Audio can be edited independently of visuals.

  • Negatives that hold: "No subtitles, no background music" finally sticks, with the right syntax (see below).

Ideal use cases

  • UGC-style and multi-register spots: One character across wildly different visual styles via text-only casting.

  • Video ads: 30s TV cut, 15s pre-roll, and 8s paid social from one master script.

  • Localized campaigns: Single-master, multi-market declination with native lip-sync and forced on-screen text language.

  • Product films: Staged, multi-beat narratives with consistent identity and set.

  • Surgical video edits: Swap an object, strip the music, keep everything else frame-identical.

Weaknesses

  • Human reference images trip content policies often, even innocuous ones (headshots, lookbook photos). Workarounds: describe characters in text with a [CHARACTER] block; when the exact face is mandatory, use a clean single-subject, neutral-background image and budget re-rolls.

  • Practical reference limits sit slightly below the on-paper caps: 1–8 distinct subjects from images, 1–5 from video, reference clips of 5–10s, edit sources under 20s. Past these, results can get lottery-like.


How to use effectively

Write a shot plan, not a paragraph. Structure every prompt as a timeline: [0-Xs] [action]. End state: [landing frame]. One main action per window — stack two and the model picks one, or blends both badly.

Name references by their order.

Bind each asset with its ordinal (@Image 1, @Image 2, @Video 1) matching input order.

Give every reference a role. One binding line per asset: @Image 1 defines <HostA>'s face, glasses, and denim jacket. Do not use the background. Several angles of one character means several images, never a collage.

End states carry continuity. Each beat's closing frame is what the next beat inherits. Omit it and props teleport, poses reset.

Lock identity explicitly. Close every prompt with character, wardrobe, and set layout. On long runs: same face, same hairstyle, same outfit, same body type for the entire video.

Treat windows as budgets. A cramped window gets a rushed beat. Size each window to the action; the model paces itself inside it.

The music kill switch. A plain "no music" often loses. The only directive that reliably holds: [SOUND] Strictly only naturally occurring sound and foley, no music allowed.

Write transitions as directives. → WHIP PAN RIGHT on her turn, smears to white, hard cut. Undescribed junctions are where objects appear and vanish.

Extend conservatively. Use your base video as a reference and describe only the new material: Extend the video naturally, [new content only], smooth motion continuity, no hard cuts, nothing appears out of thin air. The base footage is never regenerated. Ceiling: 60s final.

Audio syntax: (music) · <sound effects> · {dialogue} · 【subtitles】. For non-English lines, state language and accent before the dialogue.

Last updated