Blog · 2026-08-04
ByteDance shipped Seedance 2.5 with something unusual for a video model: a real prompting manual. Most of what circulates about video prompting is folklore. This one is written by the people who trained the thing, and it is specific — how many reference assets it will accept, how to address them individually, how to stage a 30-second clip so it does not collapse into mush, and how to edit one object in an existing video without re-rolling the whole shot.
This is an English walk-through of that guide. The technical facts — limits, syntax, the auto-locked parameters — are ByteDance's; the examples here are our own, written to be copy-pasteable rather than translated. Where the official guide is silent or ambiguous, we say so instead of guessing.
Status check, run today (4 August 2026): Seedance 2.5 is not callable through our API yet. We probed our production Volcengine Ark account directly — doubao-seedance-2-5-260128, doubao-seedance-2-5, and the -pro- and -lite- variants all return InvalidEndpointOrModel.NotFound, while the same request against doubao-seedance-2-0-260128 returns a task id. So the guide is out ahead of API access. Most of the technique below transfers directly to Seedance 2.0, which you can call today from $0.092/s.
Everything else in the guide is an elaboration of one line:
Subject + action or event
+ scene and environment (optional)
+ visual style (optional)
+ camera work or cutting (optional)
+ sound (optional)Each part has a job. Subject + action says who or what is doing what — summarise the whole movement first, then add detail only for the beat that matters, and never describe the same action twice. Scene covers place, time of day, weather and spatial relationships. Visual style is light, colour, material and texture. Camera is shot size, angle, movement, focus target and how shots join. Sound is dialogue, timbre, ambience, effects and music.
A worked example in that shape:
A luthier in a narrow workshop clamps a violin neck into a vice,
then lifts the finished body up to eye level against the window.
Late afternoon light rakes across sawdust in the air; varnish shows
a deep amber, tools sit in a worn leather roll.
Medium shot on the clamping, then a slow push to a close-up of the
grain, then a cut to the body held against the window.
Keep the scrape of the plane, the vice ratchet, and quiet room tone.Generation parameters — resolution, duration, aspect ratio — do not belong in the prompt. Set them on the request.
Seedance 2.5 accepts up to 50 reference assets in one job. The per-type ceilings, and the ranges ByteDance says are actually stable:
Images up to 30, each <= 4K stable: 1-8 subjects
Video up to 10 clips, <= 30s total stable: 1-5 subjects, 5-10s each
Audio up to 10 clips, <= 30s total stable: only what the shot needs
Editing source video + reference images stable: source <= 20s, 1-5 refsYou can push past the stable range — the guide allows 9-12 image subjects, 6-10 audio/video subjects, 6-8 editing references — and it warns plainly that stability degrades as you add more. One genuinely useful detail: if a subject needs several viewpoints, upload separate images per angle rather than one contact sheet. Independent views align more reliably than a grid crammed into a single frame.
This is the part people skip, and it is the part that decides whether a multi-reference job works. The mapping between an asset and the thing it defines must be written in the prompt. Do not rely on labels burned into the image, and do not leave the model to infer which of six photos is the protagonist.
@image1 -> <BAKER>: face, hair, and the flour-dusted grey apron.
Do NOT take the background.
@image2 -> <BAKERY>: counter layout, tiled wall, window position.
Do NOT take the people in the frame.
@video1 -> the rhythm of shaping dough and sliding the tray in.
Do NOT take the person, their clothing, or the room.
<BAKER> shapes a sourdough loaf in <BAKERY> and slides it into
the deck oven.Two habits worth forming. First, name your subjects and bind each name to exactly one asset. The guide is explicit that "@image1 through @image4 define four characters" is a broken instruction: it never says which is which. Second, write negatives only where an asset would otherwise bleed something in. Everything else should stay positive description.
When several images are the same object from different angles, say so, and say how many of the object should exist:
@image1, @image2, @image3 and @image4 all define the SAME folding
lamp - front, left, right and back. There is only ever one lamp
in the finished shot.And when a reference video already carries the choreography, do not re-describe it move by move. State what you are inheriting. Re-describing invites a conflict between your text and the asset, and the asset usually loses.
Plain language works, but four bracket types disambiguate when a shot has several audio layers at once:
( ) music (soft piano under the whole scene)
< > sound fx <a distant bell>
{ } dialogue {You came back.}
【 】 on-screen 【Chapter One】Sound exclusions are one of the few places negatives are the right tool — "no background music, keep dialogue, room tone and foley", or simply "no subtitles", or "no audio at all".
If dialogue is not in Chinese, name the language before the line, because the model otherwise drifts toward the language of the surrounding prompt. The guide's formula is language + regional variant or accent + delivery + speaker + {line}:
Spoken in American English. A young woman says it flatly,
almost bored: {You said that last time.}Long clips fail by trying to hold one undifferentiated description across half a minute. The fix is to break the story into consecutive stages, give each stage exactly one state change, and — the important bit — write the end state of each stage in terms you could verify by pausing the video.
[GOAL] A 30s process film. Subject: <CERAMICIST>. Event: throwing,
trimming and shelving a single bowl.
[STAGE 1]
Opens with: wet clay centred on the wheel, tools laid out to the right.
Event: <CERAMICIST> opens the clay and pulls the wall up.
Ends with: the bowl stands on the wheel; both hands leave the clay.
[STAGE 2]
Carried over: same person, same apron, same bowl on the wheel.
Event: the rim is trimmed with a wire tool.
Ends with: the trimmed bowl sits centre-wheel, wire tool on the tray.
[STAGE 3]
Event: the bowl is lifted onto the drying shelf.
Ends with: the bowl alone on the middle shelf, wheel empty and still.
[KEEP CONSTANT] identity, apron, wheel position, tool ownership,
left-right layout, room tone.Reach for timestamps only for hard beats — an entrance, an exit, a transition, a cut you actually care about. Stages handle ordinary narrative better. When you do use time, the guide distinguishes ranges (a budget for a beat), points ("at 5s the camera whips left") and relative timing ("three seconds after the button is pressed, the lights fade").
Timestamps are a budget, not an edit point. Segments should be continuous and non-overlapping, and an action may land slightly either side of a boundary. Too little content in a window invites the model to invent; too much produces frantic cutting or dropped beats. Do not use timestamps to demand three actions in one second.
Editing is where Seedance 2.5 earns its keep, and where prompt discipline matters most. The shape is always: declare one master, name the target, bound the change, and list what must not move.
[EDIT] Edit @video1. Between 4s and 7s only, change the wall light
on the right from cold blue to warm amber.
[MASTER] @video1 is the sole master: people, layout, action, framing,
camera movement, audio and event order all come from it.
[SCOPE] Only the right-hand wall and the area it lights. Skin tone may
follow the ambient change naturally.
[KEEP] Identity, clothing, expression, position, movement, room
structure, camera motion, dialogue and room tone: as in @video1.For a subject swap, add an explicit count and an inheritance clause. This is the single most useful sentence in the whole guide:
The white lamp inherits every entrance, movement, occlusion and exit
of the original yellow lamp: the same timings, durations, paths and
changes of speed. There is exactly one lamp in the finished video.For a background swap, the reference supplies layout, depth, ambient colour and light direction — and explicitly not the people or foreground objects in it. Keep the subject's outline, features, wardrobe, scale and motion pinned to the master.
Audio edits work the same way and can be addressed independently: strip the music while keeping dialogue, lip sync, ambience and foley; or change one speaker's language while holding their lines and timing fixed.
Three task types override what you set on the request. This trips people up because the setting is silently ignored rather than rejected:
Video editing aspect: inherited from the input, not settable
duration: matches input (+/- ~0.3s from frame handling)
First/last frame aspect: taken from the first-frame image
duration: settable
Video extension aspect: inherited from the input, not settable
duration: settableIf your first and last frame images have different aspect ratios, the last frame gets stretched. Match them.
Extension generates beyond an existing boundary, so the boundary frame is the whole game. Extending forward, describe continuity with the source's last frame first, then what happens next. Extending backward, describe the new material first and then write the source's first frame as an explicit end state — otherwise the model tends to arrive early and then keep changing the picture, or drag later characters into the earlier footage.
@video1 is the source to extend forward.
The first frame of the extension continues directly from the last
frame of @video1: same locked-off medium shot, the kite in the same
position and heading, same dune ridge behind, same late light, same
wind noise, same rightward drift.
Then the kite continues right and leaves frame; the marram grass
keeps moving in the same wind.
Throughout: same identity and clothing, same props, same background
layout, same camera axis, same ambience. One continuous subject -
no duplication, no splitting, no changing part counts.That last sentence is boilerplate worth keeping. Duplicated or splitting subjects are the characteristic failure of extension jobs.
In multimodal reference mode you can declare first and last frames inline — no separate mode needed — and still use other images for identity and materials. Declare each anchor separately; "images 1 and 2 as the first and last frames" is too vague to bind.
For multi-stage control, open with "use @image1 through @imageN as keyframes in order", then say what visible state each one represents. Independent images align better than a grid. Understand what this controls: stage order and key states, not frame-accurate reproduction.
Storyboard grids give overall story, shot order and rough composition. Keep them under about 15 cells, use clean line art, minimise text labels, and state the reading order explicitly. Then describe each shot's action and framing in the prompt, plus the final look and sound — because the grid's own line-art style is exactly what you do not want inherited.
White-box references split in two, and picking the wrong one wastes the asset:
Coarse simple geometry standing in for blocking, paths, camera and
cuts. Map EVERY solid to a final subject, and supply looks
from separate images.
Fine a finished model that only needs re-rendering. Keep structure,
motion and camera; specify materials, colour, cast and style.
Clean the plate first - no trajectory lines, axes or camera
frustums."Tense", "warm", "oppressive" set a direction, and the model will fill in the performance however it likes. To control acting, name what a viewer could actually see or hear: eyes, brow, mouth corner, breath, gaze direction, hands. Pick two to four of the clearest signals for a single emotional turn — listing every facial detail does not help. Only stage it across multiple beats if the emotion genuinely turns more than once.
Standard camera vocabulary works as-is: shot sizes, push/pull/pan/track/orbit/crane, low angle, overhead, POV. Popular moves — oner, dolly zoom, aerial, FPV, bullet time, handheld, speed ramp — also work, but if there are several subjects you still have to say who the move is built around, where it starts and where it ends.
For anything niche, keep the term and then translate it into observable change. The formula is term + what it acts on + how the picture changes + foreground/background relationship + direction or speed:
Rack focus: focus moves from the foreground leaves to the figure
behind. The leaves soften; the face resolves from blurred to sharp.Aperture, focal length and shutter values are accepted, but describing the visible result is more reliable than quoting a number.
The guide closes with a checklist. Condensed, the questions that actually catch problems: is the subject and its main action stated? Does every asset say what to take and what to ignore? Is each person, product and prop named and bound to one asset? Are assets selected per scene rather than all forced on screen at once? Does each stage have exactly one change and a verifiable end state? For edits — one master, bounded scope, explicit target count, listed keeps? Are abstract emotions and niche camera terms cashed out as something observable?
Worth internalising, because these are the requests that produce disappointment: timestamps allocate pacing and are not frame-accurate cut points; edit prompts raise the odds of alignment with the source but cannot guarantee frame-exact overlap; multi-asset work is about selecting the right asset per scene, not showing all of them; and anything that must be exactly right — subtitles, formulas, signage, product specs, frame-precise timing — should be composited beforehand or fixed in post, not asked for in a prompt.
Seedance 2.5 is not open through us yet — verified against our own Ark account today, and we will not list a model whose generate button would fail. Seedance 2.0 is live from $0.092/s and shares most of the grammar above: named subjects, per-asset roles, staged events with end states, native synced audio, up to 9 reference images plus reference video and audio. The formula, the naming discipline and the stage-with-end-state structure all transfer. When 2.5 opens, the same key and endpoint reach it — you change one model string.
Source: ByteDance's official Seedance 2.5 prompting guide. Limits, syntax and auto-locked parameters are theirs; examples and phrasing here are ours. Status figures were measured on 4 August 2026 and we re-check them.
Not yet, measured on 4 August 2026. We probed our production Volcengine Ark account: every Seedance 2.5 model id returns InvalidEndpointOrModel.NotFound, while doubao-seedance-2-0-260128 returns a task id on the same request. Seedance 2.0 is live from $0.092/s and shares most of the prompting grammar; when 2.5 opens you change one model string.
Up to 50 in one job: at most 30 images (each up to 4K), 10 video clips totalling 30s or less, and 10 audio clips totalling 30s or less. ByteDance gives narrower stable ranges — roughly 1-8 image subjects and 1-5 video/audio subjects at 5-10s each — and states plainly that stability drops as you add more. For several viewpoints of one subject, separate images align better than one contact sheet.
Because the prompt did not pin them. Two clauses fix most of it: state the exact count in the finished video, and write an inheritance clause — the new object inherits every entrance, movement, occlusion and exit of the original, with the same timings, durations, paths and speed changes. Also name the source video as the sole master so scene, camera and event order are not up for negotiation.
No. ByteDance is explicit that timestamps allocate pacing rather than mark cut points, and an action may land slightly either side of a boundary. Keep segments continuous and non-overlapping. If something must be frame-exact — subtitles, product specs, a precise cut — composite it beforehand or fix it in post rather than asking the prompt for it.