I’ve been using chatbots like ChatGPT and Claude to optimize prompts for specific video models. Lately, I feel as if the models are too verbose and off-task, even when provided the official documentation. It’s frustrating as they keep making the same mistakes. I created this set of instructions and it’s helped reduce some of the chatbot blathering and over-structuring. I also compare sample prompts below. Good luck!
When writing or optimizing prompts for Seedance 2.5, follow the style and practices demonstrated in official ByteDance Seed documentation.
Seedance prompts should normally use natural descriptive prose rather than a rigid prompt template. Temporal segmentation such as timestamps is appropriate when the shot requires precise choreography, editing, camera changes, or multiple events. Do not organize the prompt into labeled sections, categories, headings, bullet points, JSON, or a predetermined sequence such as Subject → Action → Camera → Style.
There is no mandatory prompt order. Arrange information in whatever order most naturally and clearly describes the requested video.
Treat concepts such as subject, action, setting, camera, references, dialogue, visual style, sound, and timing as an internal checklist only. Include only the ones that matter to the specific shot. Never expose that checklist as the structure of the finished prompt.
Match prompt detail to shot complexity. A simple shot should produce a short prompt. A complicated sequence involving multiple characters, references, actions, locations, camera changes, or precise timing may require substantially more detail.
Before rewriting, identify and resolve internal contradictions in the source prompt. Prefer fixing conflicting instructions over adding extra safeguards. Do not preserve mutually incompatible camera movement, geography, timing, visibility, action, lighting, or audio instructions merely because the user supplied them.
Preserve the user's creative decisions. Optimization means making the request clearer for Seedance, not adding extra cinematography, performance direction, visual adjectives, safeguards, negative instructions, or production details the user did not request.
Use references naturally and explicitly when they are provided. State what a reference represents or controls where necessary, following patterns used by ByteDance such as:
`Use @Image 1 for the venue.`
`Reference @Image 2 for the pianist.`
`The lead vocalist must strictly follow @Image 5.`
`Refer to @Clay Render 1 for camera movement, pacing, shot-size transitions, subject trajectory, and blocking.`
Use references naturally and explicitly. References may be bound inline where the referenced element becomes relevant, or grouped together when several references establish the cast, environment, instruments, audience, or other elements before the action begins. Choose whichever structure makes the relationship between each reference and the requested video clearest. Do not force references into a separate block or force them inline when either approach would make the prompt less natural.
For named principal characters, bind the character clearly to the relevant reference the first time the character becomes important. Once established, use the character's name naturally rather than repeatedly rebinding the same reference.
Use stronger language such as `must strictly follow` selectively for identity-critical principal characters or references where exact fidelity matters. Do not automatically apply it to every reference.
When a reference controls only one aspect of a subject, say so directly when useful:
`@Image 6 for appearance only.`
`Refer to @Video 1 for motion and timing.`
`Use @Image 2 for wardrobe.`
Do not mechanically enumerate reference properties when the intended role is already obvious.
Do not repeatedly redescribe information already established by a reference. Restate specific appearance, wardrobe, prop, or material details only when they are important to the requested result, not reliably communicated by the reference, or have previously failed and therefore need reinforcement.
Describe events chronologically when chronology matters, but do not artificially convert a shot into stages or beats.
Use timestamps only when specific timing meaningfully improves control, such as complex choreography, several distinct events, timed camera changes, transformations, multi-shot sequences, or targeted video edits. Do not timestamp a simple shot merely because its duration is known.
When timestamps are useful, integrate them directly into the prose:
`0–5s: ... 6–10s: ... 11–20s: ...`
Do not turn each timestamp into a separate formatted section unless the user requests that format.
For continuous shots, state that the shot is continuous when important and then describe how the action and camera evolve naturally. Do not add repeated continuity safeguards throughout the prompt.
For dialogue, place the spoken words directly into the action where they occur. Add performance direction only when it materially affects the intended performance.
For audio references, bind the audio clearly once when possible:
`Use @Audio1 as Junior's continuous dialogue track for the full 15 seconds.`
Do not repeatedly restate the same audio binding inside every timed segment unless the relationship changes. Let the shot description establish when the speaker is on-camera, off-camera, close, distant, lip-synced, or otherwise changes spatial relationship to the audio.
For video extension, follow the existing video rather than rebuilding its description. Identify the source video, specify only the continuity that matters, and describe what happens next.
Example:
`Extend @Video 1. Continue from the visuals and subjects in @Video 1, keeping the characters, scene, visual style, and sound consistent. [What happens next.]`
For video editing, identify the source, state what must remain unchanged, identify the target change, and describe that change.
Example:
`Edit @Video 1. Keep the characters, actions, and visual style unchanged. Adjust only the camera movement. [New camera behavior.]`
For reference-to-video, tell Seedance what useful information to take from each source when that distinction matters. Do not mechanically list every property a reference might control.
For multi-character scenes, introduce character references inline with the moment those characters appear. If spatial separation, seating, blocking, or ordering is important, state it clearly enough to avoid ambiguity. Do not add generic multi-character safeguards unless a specific problem requires them.
Do not begin prompts with labels such as `R2V prompt:`, `T2V prompt:`, `Prompt:`, or `Seedance 2.5 Prompt:` unless the user explicitly wants the label.
Do not automatically append generic instructions such as:
* maintain identity consistency
* no morphing
* no warping
* realistic physics
* correct anatomy
* stable faces
* stable lighting
* no artifacts
* preserve wardrobe
* cinematic quality
* smooth natural motion
Add a constraint only when it is relevant to the user's request, resolves a known failure, or performs a specific function in a reference, extension, or editing workflow.
When a character or object must not occlude another subject, when lighting must change at a specific threshold, or when spatial geography is essential, state that relationship directly rather than relying on generic preservation language.
Do not embellish a prompt merely because more detail is possible. Avoid redundant adjectives, duplicated continuity instructions, repeated reference bindings, and restating information that Seedance already has from the supplied assets.
When the user's original wording already communicates something effectively, retain it rather than translating it into more elaborate "prompt language."
When the user requests a sparse prompt or maximum model creativity, reduce the prompt to the minimum information needed to preserve the intended result. Keep required characters and references, essential actions, critical spatial relationships, dialogue or audio, required cuts or transitions, and any camera, timing, or continuity instructions that materially define the shot.
Remove details that are optional rather than automatically removing entire categories of information. Omit unnecessary staging, micro-actions, descriptive embellishment, redundant reference information, and camera or timing instructions that do not affect the intended result.
Sparse does not mean vague. Keep any detail whose omission could materially change the scene, performance, sequence of events, spatial logic, or camera behavior the user actually cares about. Leave everything else open for Seedance to interpret.
When shortening a prompt to meet a platform character limit, compress in this order:
1. Remove duplicated instructions.
2. Remove repeated reference descriptions.
3. Combine sentences describing the same action or camera move.
4. Remove nonessential adjectives and atmosphere.
5. Shorten character descriptions already represented by references.
6. Preserve essential references, actions, timing, geography, dialogue/audio, cuts, and known failure corrections.
Do not solve character-limit problems by stripping away the core shot logic.
The finished prompt should feel like a director plainly describing the desired video to Seedance, not like a prompt engineer filling out a template.
Use official ByteDance Seed examples as the primary style reference. Official Dreamina, CapCut, and Volcengine documentation may supplement them. Do not use third-party prompting conventions as Seedance best practices.
When official documentation does not establish that a technique is beneficial, do not present it as an optimization. If making an inference from documented examples or observed model behavior, distinguish it from documented guidance.
Most importantly: do not optimize for apparent sophistication. Optimize for fidelity to the user's intended shot using the minimum amount of clear instruction necessary.
Why I made the Seedance 2.5 optimizer
I gave both ChatGPT and Claude the Seed blog about prompting with references for their great examples.
T2V prompt: One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.
R2V prompt: 16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot from @Image 1, following her water sleeves through the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side in a combat exchange. Consort Yu stands slightly behind and to the side of the Overlord, weaving in water-sleeve movements to set softness against strength. The camera slowly pulls back from a medium-close shot of the warrior to a full stage view. At the end, all three face the audience and strike a synchronized Peking opera finale pose.
R2V prompt: A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting and shadows, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly follow @Image 5. Reference @Images 6 to 10 for the rest of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the audience seating. The lead vocalist walks from center stage toward the front edge. The pianist is positioned by the piano. The orchestra is arranged on both sides and toward the rear. The choir stands at the back of the stage. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera naturally moves across the violin, cello, and orchestra as they perform together, with the violin feeling bright and the cello warm. In the latter part, the choir joins in. The lead vocalist briefly makes eye contact with front-row audience members, who respond with a smile and a slight nod. In the closing shot, the camera pulls back. The singing ends, and the audience joins in the applause.
Those all look simple, readable, and cinematic. But both Claude and ChatGPT started giving me very lengthy and highly categorized prompts, which suggests the output of how an LLM thinks more than an optimized Seedance prompt. Note all of the headers and categorization that is not present in the Seed examples.
30-second cinematic realistic video, 16:9 landscape.
ONE SINGLE CONTINUOUS TAKE for the entire 30 seconds. No cuts, transitions or time jumps.
VISUAL AND CAMERA LANGUAGE: Cinematic realism with a natural 35mm cinema-lens perspective. Smooth, controlled Steadicam-style photography with the physical weight and inertia of a real camera operator. Moderate depth of field, natural focus falloff, subtle cinematic motion blur, realistic skin and material textures, soft highlight roll-off and restrained contrast. Warm tungsten practical lighting throughout the hotel contrasts naturally with cool blue late-night window light. The fundamental camera movement is a CONTINUOUS BACKWARDS TRACKING SHOT. Mike walks forward through the hotel while facing generally toward camera, and the camera continuously retreats ahead of him at his walking pace. Mike's face remains clearly visible throughout the shot. The camera never passes behind Mike and never allows him to overtake it. Reframings toward other characters happen while the camera continues travelling backwards.
LOCATION AND GEOGRAPHY: Late night inside the lobby of a grand old hotel as a gala winds down. Marble floors, brass fittings, chandeliers, table lamps, deep carpet and tall windows. The lobby is half-emptied and quiet. The lobby is one long straight hall. Mike enters at the near end and walks its full length toward the far end without stopping. From Mike's direction of travel, a long illuminated bar runs along his LEFT side and a seating lounge of armchairs runs along his RIGHT side. At the far end is a low stage with a covered grand piano. Just before the stage, beside the elevator bank, a service corridor opens from Mike's RIGHT. Maintain this physical geography consistently for the entire take.
Mike must strictly follow @image1 he leads the shot and is on screen throughout.
Reference @image2 for Devin.
Reference @image3 for Frogman.
Reference @image4 for Hacker.
Reference @image5 for Cardy.
Reference @image6 for Guy.
Reference @image7 for Busker.
Reference @image8 for Lava Lamp Torso Woman with the glowing amber panel in her dress.
Reference @image9 for Mechanic.
Reference @10 for Harper.
Use @Images 6, 7 and 8 for appearance only; do not carry over their original backgrounds or lighting.
No other people are present anywhere in the lobby.
0–3s: Begin just inside the hotel entrance facing the brass-framed glass door. Devin holds it open as Mike steps in from the rain and walks toward camera. Devin subtly tips his head as Mike passes. The camera begins retreating with Mike, the rainy entrance receding behind him.
3–7s: The camera gently arcs to reveal the lounge while keeping Mike prominent. Frogman sits at its near end reading a folded newspaper. As Mike passes, Frogman looks up over the newspaper at him, then returns to reading.
7–12s: The camera smoothly arcs across to reveal the bar while keeping Mike prominent. Cardy stands behind it polishing a glass. Guy sits at the near end facing the hall and tracks Mike only with his eyes, without turning his head. Brief rack focus from Mike's face to Guy's eyes, then back to Mike.
12–17s: The camera arcs back toward the lounge. Further down, Hacker is folded into a wingback chair with a laptop, blue screen light on her face; Mechanic sits opposite with elbows on his knees. Ahead, the woman with the Lava Lamp Torso Woman crosses from the lounge side to the bar side, the lava slowly moving in her amber torso as she crosses Mike's path.
17–23s: The camera recentres and gradually closes the distance to Mike, tightening into a moving portrait without letting him overtake it. As its retreat carries it alongside the low stage, the camera eases off Mike's line so the stage comes into frame beside him. Busker sits on its edge beside the covered grand piano, playing quietly. She glances up at Mike; Mike briefly looks toward her as he passes.
23–27s: Mike turns RIGHT into the service corridor. The camera anticipates his turn, curving smoothly backwards into the corridor mouth ahead of him while maintaining a frontal view. Harper has already emerged into the lobby pushing a housekeeping cart. Mike passes behind the moving cart, then follows the camera into the corridor.
27–30s: The camera retreats ahead of Mike through the bare fluorescent service corridor as the warm lobby falls away behind him. Camera and Mike decelerate to a stop. Hold briefly on his face as the gala ambience dies away. After a beat, Mike gives an understated shrug as if asking the camera operator if the shot worked.
AUDIO: Natural synchronized sound only. No score and no dialogue. Rain and faint street noise at the entrance fading into lobby hush; distant piano; occasional ice in a shaker; one elevator chime. Busker's quiet playing becomes more audible as Mike approaches and fades as he passes. Lobby ambience recedes upon entering the service corridor, ending in sparse fluorescent corridor room tone.
Here is an example when I pushed back and said “Are you sure this is optimized for Seedance 2.5?”
A 30-second cinematic realistic video in 16:9 landscape, shot as one continuous take with no cuts or time jumps. Late night in the lobby of a grand old hotel as a gala winds down — marble floors, brass fittings, chandeliers, table lamps, deep carpet and tall windows, the room half-emptied and quiet. Warm tungsten practicals throughout the lobby against cool blue late-night window light. A 35mm cinema-lens perspective with smooth Steadicam-style movement that carries the weight and inertia of a real operator. The camera remains ahead of Mike throughout, retreating as he walks and adjusting its speed and lateral position for each reframe. Mike, who must strictly follow Mike @image1, walks forward down the hall facing generally toward camera. He is relaxed and quietly amused by the strange people around him, looking naturally from side to side as he notices them, with small knowing smiles and understated reactions rather than walking blankly forward. The lobby is one long straight hall. Mike enters at the near end and walks nearly its full length before turning right into a service corridor just before the stage. From his direction of travel, a long illuminated bar runs along his left and a seating lounge with several separate armchairs runs along his right. At the far end is a low stage with a grand piano, and just before the stage, beside the elevator bank, a service corridor opens on his right. No other people beyond the referenced characters are present. 0–3s: Begin just inside the entrance facing the brass-framed glass door. Mike is outside in cool blue rainy night light with wet street reflections. Devin, following @image2, holds the door open as Mike steps inside. Warm lobby light gradually takes over only after Mike crosses the threshold. Mike takes in the room with a faintly amused look as he walks toward camera. Devin follows him inside and closes the door. 3–7s: The camera arcs gently toward the lounge while keeping Mike prominent. Frogman, following @image3, sits alone in one armchair wearing his tank top and cut-off jeans, reading a folded newspaper. Hacker, following @image4 sits in a separate armchair nearby with a laptop, blue screen light on her face. An empty third chair sits beside them. Devin walks behind Mike to that empty chair and sits in it, clearly separate from Frogman and Hacker. Mike glances toward Frogman with restrained amusement. Frogman looks up over the newspaper at Mike, then returns to reading. 7–12s: The camera arcs across to the illuminated bar and clearly reveals Cardy, following @image5, an android standing behind the bar in a red shirt, polishing a glass. Cardy remains clearly visible as Mike passes. Guy, @image6 for appearance only, sits at the near end of the bar facing the hall and tracks Mike with his eyes without turning his head. Mike notices Cardy, then catches Guy watching him and gives the slightest amused reaction. Brief rack focus from Mike to Guy's eyes, then back to Mike. 12–17s: The camera arcs back toward the lounge. Devin remains seated in his own chair beside Frogman and Hacker. Mechanic, following @image7, sits opposite with his elbows on his knees. Lava Lamp Torso Woman, referencing @image8 for appearance only, crosses from the lounge side to the bar side without passing in front of or occluding Mike but waving to him. The lava moves slowly inside the glowing amber panel of her dress. Mike watches her cross with a small incredulous smile, then looks ahead again. 17–23s: The camera recentres and slightly slows its retreat, tightening into a moving portrait of Mike. As the stage comes into frame beside him, Busker, referencing @image9 for appearance only, sits at the grand piano playing quietly. She glances up at Mike. Mike looks toward her, gives a subtle appreciative half-smile, then continues walking. 23–27s: Mike turns right into the service corridor. The camera anticipates the turn, curving smoothly backwards into the corridor mouth ahead of him while staying frontal. An elevator chime sounds as Harper, following @image10, emerges pushing a housekeeping cart, wearing a black tank top and camo cargo shorts and makes eye contact with the camera. Harper and the cart stay clear of Mike's path and never occlude him. Mike gives her a quick amused glance as he passes beside the cart and enters the corridor. 27–30s: The camera retreats ahead of Mike through the bare fluorescent service corridor as the warm lobby falls away behind him. Mike looks back briefly toward the bizarre lobby he has just crossed, then back toward camera with quiet self-amusement. Camera and Mike decelerate to a stop. After a beat, Mike gives an understated shrug, as if asking the camera operator whether the shot worked. Natural synchronized sound only, no score and no dialogue. Rain and faint street noise at the entrance fade into the lobby hush. Busker's quiet piano grows more audible as Mike approaches the stage and fades as he passes. The elevator chime sounds just before Harper emerges. The lobby ambience recedes as Mike enters the service corridor, ending in sparse fluorescent room tone.
I realized it wasn’t really taking my idea and optimizing it, but rather optimizing the non-optimized version. This means a lot of clutter was left in the prompt.
This was when I decided to build instructions (listed at the top). Here is the prompt I received. Notice how much more readable this is (aside from the constant “must strictly follow” instruction).
A cinematic 30-second single continuous shot in late afternoon on a lively Paris street with narrow sidewalks, Haussmann-style stone façades, café terraces, parked scooters, and market stalls. Mike must strictly follow @image1. Filmed on an ARRI Alexa as a single continuous Steadicam shot, with the operator walking backward in front of Mike and maintaining a natural human-operated feel through subtle footstep drift, gentle lateral corrections, and organic reframing. Mike can move off-center as foreground pedestrians, café furniture, and market stalls create natural depth and occasional brief occlusion, with the framing breathing between medium-wide and medium shots without obvious zooms. Mike walks down the street with AlbumGuy, who must strictly follow @image2, on his left and Lava Lamp Torso, who must strictly follow @image3, on his right. The lava inside her torso visibly moves slowly and continuously. The three are walking together. Mike is relaxed and quietly amused by the people he encounters, looking naturally from side to side as he notices them, with small knowing smiles and understated reactions rather than walking blankly forward. Mike first passes Rory, who must strictly follow @image4, leaning against a brick wall with a hole in it. Rory gestures to the hole and shrugs. He then passes Eleanor, who must strictly follow @image5 and Android who must strictly follow @image6 who are deep in an emotionally moving conversation. Mike passes Devin, who must strictly follow
@image7, and his father, who must strictly follow @image8, smiling together at a market stall looking at a fish. Mike passes Busker, who must strictly follow @image9, is playing guitar. At this point, AlbumGuy stops to listen to the Busker while Mike continues walking with Lava Lamp Torso. Mike then passes Frogman, who must strictly follow @image10, who is sipping a coffee with a croissant on a plate in front of him. At this point, Lava Lamp Torso with Frogman and takes a bite of the croissant. Mike continues forward alone with a shrug. Hacker, who must strictly follow @image11, bumps into Mike and reaches into his back pocket to steal his wallet. She then walks away holding the wallet up and looks back toward the camera with a smirk. After the wallet is taken, the camera swings from the front-facing view into a profile view of Mike as Hacker walks away. This profile reframing reveals the adjacent intersecting street and creates room in the composition for the new street to open up beside Mike. Mike is joined on the left side by Concierge, who must strictly follow @image12 in her blue suit, and on the right side by Harper, who must strictly follow @image13, in her blank tank top and camo cargo shorts, both walking with Mike. Harper playfully bumps into Mike like old friends reunited. As they reach an intersection, Concierge gestures toward the Eiffel Tower visible in the distance down the intersecting street. Each referenced character appears only during their described encounter. Referenced characters must not appear as background pedestrians before their encounter and must not reappear after Mike has passed them, except AlbumGuy and Lava Lamp Torso, who begin with Mike and remain until they stop at their respective encounters. The camera gives each described character or pair a clear readable moment before continuing forward. No music, only ambient Paris street sound.












