Author: Mike

  • Seedance 2.5 Prompt Optimizer

    I’ve been using chatbots like ChatGPT and Claude to optimize prompts for specific video models. Lately, I feel as if the models are too verbose and off-task, even when provided the official documentation. It’s frustrating as they keep making the same mistakes. I created this set of instructions and it’s helped reduce some of the chatbot blathering and over-structuring. I also compare sample prompts below. Good luck!

    When writing or optimizing prompts for Seedance 2.5, follow the style and practices demonstrated in official ByteDance Seed documentation.
    
    Seedance prompts should normally use natural descriptive prose rather than a rigid prompt template. Temporal segmentation such as timestamps is appropriate when the shot requires precise choreography, editing, camera changes, or multiple events. Do not organize the prompt into labeled sections, categories, headings, bullet points, JSON, or a predetermined sequence such as Subject → Action → Camera → Style.
    
    There is no mandatory prompt order. Arrange information in whatever order most naturally and clearly describes the requested video.
    
    Treat concepts such as subject, action, setting, camera, references, dialogue, visual style, sound, and timing as an internal checklist only. Include only the ones that matter to the specific shot. Never expose that checklist as the structure of the finished prompt.
    
    Match prompt detail to shot complexity. A simple shot should produce a short prompt. A complicated sequence involving multiple characters, references, actions, locations, camera changes, or precise timing may require substantially more detail.
    
    Before rewriting, identify and resolve internal contradictions in the source prompt. Prefer fixing conflicting instructions over adding extra safeguards. Do not preserve mutually incompatible camera movement, geography, timing, visibility, action, lighting, or audio instructions merely because the user supplied them.
    
    Preserve the user's creative decisions. Optimization means making the request clearer for Seedance, not adding extra cinematography, performance direction, visual adjectives, safeguards, negative instructions, or production details the user did not request.
    
    Use references naturally and explicitly when they are provided. State what a reference represents or controls where necessary, following patterns used by ByteDance such as:
    
    `Use @Image 1 for the venue.`
    
    `Reference @Image 2 for the pianist.`
    
    `The lead vocalist must strictly follow @Image 5.`
    
    `Refer to @Clay Render 1 for camera movement, pacing, shot-size transitions, subject trajectory, and blocking.`
    
    Use references naturally and explicitly. References may be bound inline where the referenced element becomes relevant, or grouped together when several references establish the cast, environment, instruments, audience, or other elements before the action begins. Choose whichever structure makes the relationship between each reference and the requested video clearest. Do not force references into a separate block or force them inline when either approach would make the prompt less natural.
    
    For named principal characters, bind the character clearly to the relevant reference the first time the character becomes important. Once established, use the character's name naturally rather than repeatedly rebinding the same reference.
    
    Use stronger language such as `must strictly follow` selectively for identity-critical principal characters or references where exact fidelity matters. Do not automatically apply it to every reference.
    
    When a reference controls only one aspect of a subject, say so directly when useful:
    
    `@Image 6 for appearance only.`
    
    `Refer to @Video 1 for motion and timing.`
    
    `Use @Image 2 for wardrobe.`
    
    Do not mechanically enumerate reference properties when the intended role is already obvious.
    
    Do not repeatedly redescribe information already established by a reference. Restate specific appearance, wardrobe, prop, or material details only when they are important to the requested result, not reliably communicated by the reference, or have previously failed and therefore need reinforcement.
    
    Describe events chronologically when chronology matters, but do not artificially convert a shot into stages or beats.
    
    Use timestamps only when specific timing meaningfully improves control, such as complex choreography, several distinct events, timed camera changes, transformations, multi-shot sequences, or targeted video edits. Do not timestamp a simple shot merely because its duration is known.
    
    When timestamps are useful, integrate them directly into the prose:
    
    `0–5s: ... 6–10s: ... 11–20s: ...`
    
    Do not turn each timestamp into a separate formatted section unless the user requests that format.
    
    For continuous shots, state that the shot is continuous when important and then describe how the action and camera evolve naturally. Do not add repeated continuity safeguards throughout the prompt.
    
    For dialogue, place the spoken words directly into the action where they occur. Add performance direction only when it materially affects the intended performance.
    
    For audio references, bind the audio clearly once when possible:
    
    `Use @Audio1 as Junior's continuous dialogue track for the full 15 seconds.`
    
    Do not repeatedly restate the same audio binding inside every timed segment unless the relationship changes. Let the shot description establish when the speaker is on-camera, off-camera, close, distant, lip-synced, or otherwise changes spatial relationship to the audio.
    
    For video extension, follow the existing video rather than rebuilding its description. Identify the source video, specify only the continuity that matters, and describe what happens next.
    
    Example:
    
    `Extend @Video 1. Continue from the visuals and subjects in @Video 1, keeping the characters, scene, visual style, and sound consistent. [What happens next.]`
    
    For video editing, identify the source, state what must remain unchanged, identify the target change, and describe that change.
    
    Example:
    
    `Edit @Video 1. Keep the characters, actions, and visual style unchanged. Adjust only the camera movement. [New camera behavior.]`
    
    For reference-to-video, tell Seedance what useful information to take from each source when that distinction matters. Do not mechanically list every property a reference might control.
    
    For multi-character scenes, introduce character references inline with the moment those characters appear. If spatial separation, seating, blocking, or ordering is important, state it clearly enough to avoid ambiguity. Do not add generic multi-character safeguards unless a specific problem requires them.
    
    Do not begin prompts with labels such as `R2V prompt:`, `T2V prompt:`, `Prompt:`, or `Seedance 2.5 Prompt:` unless the user explicitly wants the label.
    
    Do not automatically append generic instructions such as:
    
    * maintain identity consistency
    * no morphing
    * no warping
    * realistic physics
    * correct anatomy
    * stable faces
    * stable lighting
    * no artifacts
    * preserve wardrobe
    * cinematic quality
    * smooth natural motion
    
    Add a constraint only when it is relevant to the user's request, resolves a known failure, or performs a specific function in a reference, extension, or editing workflow.
    
    When a character or object must not occlude another subject, when lighting must change at a specific threshold, or when spatial geography is essential, state that relationship directly rather than relying on generic preservation language.
    
    Do not embellish a prompt merely because more detail is possible. Avoid redundant adjectives, duplicated continuity instructions, repeated reference bindings, and restating information that Seedance already has from the supplied assets.
    
    When the user's original wording already communicates something effectively, retain it rather than translating it into more elaborate "prompt language."
    
    When the user requests a sparse prompt or maximum model creativity, reduce the prompt to the minimum information needed to preserve the intended result. Keep required characters and references, essential actions, critical spatial relationships, dialogue or audio, required cuts or transitions, and any camera, timing, or continuity instructions that materially define the shot.
    
    Remove details that are optional rather than automatically removing entire categories of information. Omit unnecessary staging, micro-actions, descriptive embellishment, redundant reference information, and camera or timing instructions that do not affect the intended result.
    
    Sparse does not mean vague. Keep any detail whose omission could materially change the scene, performance, sequence of events, spatial logic, or camera behavior the user actually cares about. Leave everything else open for Seedance to interpret.
    
    When shortening a prompt to meet a platform character limit, compress in this order:
    
    1. Remove duplicated instructions.
    2. Remove repeated reference descriptions.
    3. Combine sentences describing the same action or camera move.
    4. Remove nonessential adjectives and atmosphere.
    5. Shorten character descriptions already represented by references.
    6. Preserve essential references, actions, timing, geography, dialogue/audio, cuts, and known failure corrections.
    
    Do not solve character-limit problems by stripping away the core shot logic.
    
    The finished prompt should feel like a director plainly describing the desired video to Seedance, not like a prompt engineer filling out a template.
    
    Use official ByteDance Seed examples as the primary style reference. Official Dreamina, CapCut, and Volcengine documentation may supplement them. Do not use third-party prompting conventions as Seedance best practices.
    
    When official documentation does not establish that a technique is beneficial, do not present it as an optimization. If making an inference from documented examples or observed model behavior, distinguish it from documented guidance.
    
    Most importantly: do not optimize for apparent sophistication. Optimize for fidelity to the user's intended shot using the minimum amount of clear instruction necessary.

    Why I made the Seedance 2.5 optimizer

    I gave both ChatGPT and Claude the Seed blog about prompting with references for their great examples.

    T2V prompt: One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.
    R2V prompt: 16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot from @Image 1, following her water sleeves through the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side in a combat exchange. Consort Yu stands slightly behind and to the side of the Overlord, weaving in water-sleeve movements to set softness against strength. The camera slowly pulls back from a medium-close shot of the warrior to a full stage view. At the end, all three face the audience and strike a synchronized Peking opera finale pose.
    R2V prompt: A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting and shadows, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly follow @Image 5. Reference @Images 6 to 10 for the rest of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the audience seating. The lead vocalist walks from center stage toward the front edge. The pianist is positioned by the piano. The orchestra is arranged on both sides and toward the rear. The choir stands at the back of the stage. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera naturally moves across the violin, cello, and orchestra as they perform together, with the violin feeling bright and the cello warm. In the latter part, the choir joins in. The lead vocalist briefly makes eye contact with front-row audience members, who respond with a smile and a slight nod. In the closing shot, the camera pulls back. The singing ends, and the audience joins in the applause.

    Those all look simple, readable, and cinematic. But both Claude and ChatGPT started giving me very lengthy and highly categorized prompts, which suggests the output of how an LLM thinks more than an optimized Seedance prompt. Note all of the headers and categorization that is not present in the Seed examples.

    30-second cinematic realistic video, 16:9 landscape. 
    
    ONE SINGLE CONTINUOUS TAKE for the entire 30 seconds. No cuts, transitions or time jumps. 
    
    VISUAL AND CAMERA LANGUAGE: Cinematic realism with a natural 35mm cinema-lens perspective. Smooth, controlled Steadicam-style photography with the physical weight and inertia of a real camera operator. Moderate depth of field, natural focus falloff, subtle cinematic motion blur, realistic skin and material textures, soft highlight roll-off and restrained contrast. Warm tungsten practical lighting throughout the hotel contrasts naturally with cool blue late-night window light. The fundamental camera movement is a CONTINUOUS BACKWARDS TRACKING SHOT. Mike walks forward through the hotel while facing generally toward camera, and the camera continuously retreats ahead of him at his walking pace. Mike's face remains clearly visible throughout the shot. The camera never passes behind Mike and never allows him to overtake it. Reframings toward other characters happen while the camera continues travelling backwards. 
    
    
    LOCATION AND GEOGRAPHY: Late night inside the lobby of a grand old hotel as a gala winds down. Marble floors, brass fittings, chandeliers, table lamps, deep carpet and tall windows. The lobby is half-emptied and quiet. The lobby is one long straight hall. Mike enters at the near end and walks its full length toward the far end without stopping. From Mike's direction of travel, a long illuminated bar runs along his LEFT side and a seating lounge of armchairs runs along his RIGHT side. At the far end is a low stage with a covered grand piano. Just before the stage, beside the elevator bank, a service corridor opens from Mike's RIGHT. Maintain this physical geography consistently for the entire take.
    
    
    Mike must strictly follow @image1 he leads the shot and is on screen throughout. 
    
    Reference @image2 for Devin. 
    
    Reference @image3 for Frogman.
    
    Reference @image4 for Hacker. 
    
    Reference @image5 for Cardy. 
    
    Reference @image6 for Guy. 
    
    Reference @image7 for Busker. 
    
    Reference @image8 for Lava Lamp Torso Woman with  the glowing amber panel in her dress.
    Reference @image9 for Mechanic. 
    
    Reference @10 for Harper. 
    
    Use @Images 6, 7 and 8 for appearance only; do not carry over their original backgrounds or lighting.
    
    No other people are present anywhere in the lobby. 
    
    
    0–3s: Begin just inside the hotel entrance facing the brass-framed glass door. Devin holds it open as Mike steps in from the rain and walks toward camera. Devin subtly tips his head as Mike passes. The camera begins retreating with Mike, the rainy entrance receding behind him. 
    
    
    3–7s: The camera gently arcs to reveal the lounge while keeping Mike prominent. Frogman sits at its near end reading a folded newspaper. As Mike passes, Frogman looks up over the newspaper at him, then returns to reading. 
    
    
    7–12s: The camera smoothly arcs across to reveal the bar while keeping Mike prominent. Cardy stands behind it polishing a glass. Guy sits at the near end facing the hall and tracks Mike only with his eyes, without turning his head. Brief rack focus from Mike's face to Guy's eyes, then back to Mike. 
    
    
    12–17s: The camera arcs back toward the lounge. Further down, Hacker is folded into a wingback chair with a laptop, blue screen light on her face; Mechanic sits opposite with elbows on his knees. Ahead, the woman with the Lava Lamp Torso Woman crosses from the lounge side to the bar side, the lava slowly moving in her amber torso as she crosses Mike's path. 
    
    
    17–23s: The camera recentres and gradually closes the distance to Mike, tightening into a moving portrait without letting him overtake it. As its retreat carries it alongside the low stage, the camera eases off Mike's line so the stage comes into frame beside him. Busker sits on its edge beside the covered grand piano, playing quietly. She glances up at Mike; Mike briefly looks toward her as he passes.
    
    
    23–27s: Mike turns RIGHT into the service corridor. The camera anticipates his turn, curving smoothly backwards into the corridor mouth ahead of him while maintaining a frontal view. Harper has already emerged into the lobby pushing a housekeeping cart. Mike passes behind the moving cart, then follows the camera into the corridor.
    
    
    27–30s: The camera retreats ahead of Mike through the bare fluorescent service corridor as the warm lobby falls away behind him. Camera and Mike decelerate to a stop. Hold briefly on his face as the gala ambience dies away. After a beat, Mike gives an understated shrug as if asking the camera operator if the shot worked.
    
    
    AUDIO: Natural synchronized sound only. No score and no dialogue. Rain and faint street noise at the entrance fading into lobby hush; distant piano; occasional ice in a shaker; one elevator chime. Busker's quiet playing becomes more audible as Mike approaches and fades as he passes. Lobby ambience recedes upon entering the service corridor, ending in sparse fluorescent corridor room tone.

    Here is an example when I pushed back and said “Are you sure this is optimized for Seedance 2.5?”

    A 30-second cinematic realistic video in 16:9 landscape, shot as one continuous take with no cuts or time jumps. Late night in the lobby of a grand old hotel as a gala winds down — marble floors, brass fittings, chandeliers, table lamps, deep carpet and tall windows, the room half-emptied and quiet. Warm tungsten practicals throughout the lobby against cool blue late-night window light. A 35mm cinema-lens perspective with smooth Steadicam-style movement that carries the weight and inertia of a real operator. The camera remains ahead of Mike throughout, retreating as he walks and adjusting its speed and lateral position for each reframe. Mike, who must strictly follow Mike @image1, walks forward down the hall facing generally toward camera. He is relaxed and quietly amused by the strange people around him, looking naturally from side to side as he notices them, with small knowing smiles and understated reactions rather than walking blankly forward. The lobby is one long straight hall. Mike enters at the near end and walks nearly its full length before turning right into a service corridor just before the stage. From his direction of travel, a long illuminated bar runs along his left and a seating lounge with several separate armchairs runs along his right. At the far end is a low stage with a grand piano, and just before the stage, beside the elevator bank, a service corridor opens on his right. No other people beyond the referenced characters are present. 0–3s: Begin just inside the entrance facing the brass-framed glass door. Mike is outside in cool blue rainy night light with wet street reflections. Devin, following @image2, holds the door open as Mike steps inside. Warm lobby light gradually takes over only after Mike crosses the threshold. Mike takes in the room with a faintly amused look as he walks toward camera. Devin follows him inside and closes the door. 3–7s: The camera arcs gently toward the lounge while keeping Mike prominent. Frogman, following @image3, sits alone in one armchair wearing his tank top and cut-off jeans, reading a folded newspaper. Hacker, following @image4 sits in a separate armchair nearby with a laptop, blue screen light on her face. An empty third chair sits beside them. Devin walks behind Mike to that empty chair and sits in it, clearly separate from Frogman and Hacker. Mike glances toward Frogman with restrained amusement. Frogman looks up over the newspaper at Mike, then returns to reading. 7–12s: The camera arcs across to the illuminated bar and clearly reveals Cardy, following @image5, an android standing behind the bar in a red shirt, polishing a glass. Cardy remains clearly visible as Mike passes. Guy, @image6 for appearance only, sits at the near end of the bar facing the hall and tracks Mike with his eyes without turning his head. Mike notices Cardy, then catches Guy watching him and gives the slightest amused reaction. Brief rack focus from Mike to Guy's eyes, then back to Mike. 12–17s: The camera arcs back toward the lounge. Devin remains seated in his own chair beside Frogman and Hacker. Mechanic, following @image7, sits opposite with his elbows on his knees. Lava Lamp Torso Woman, referencing @image8 for appearance only, crosses from the lounge side to the bar side without passing in front of or occluding Mike but waving to him. The lava moves slowly inside the glowing amber panel of her dress. Mike watches her cross with a small incredulous smile, then looks ahead again. 17–23s: The camera recentres and slightly slows its retreat, tightening into a moving portrait of Mike. As the stage comes into frame beside him, Busker, referencing @image9 for appearance only, sits at the grand piano playing quietly. She glances up at Mike. Mike looks toward her, gives a subtle appreciative half-smile, then continues walking. 23–27s: Mike turns right into the service corridor. The camera anticipates the turn, curving smoothly backwards into the corridor mouth ahead of him while staying frontal. An elevator chime sounds as Harper, following @image10, emerges pushing a housekeeping cart, wearing a black tank top and camo cargo shorts and makes eye contact with the camera. Harper and the cart stay clear of Mike's path and never occlude him. Mike gives her a quick amused glance as he passes beside the cart and enters the corridor. 27–30s: The camera retreats ahead of Mike through the bare fluorescent service corridor as the warm lobby falls away behind him. Mike looks back briefly toward the bizarre lobby he has just crossed, then back toward camera with quiet self-amusement. Camera and Mike decelerate to a stop. After a beat, Mike gives an understated shrug, as if asking the camera operator whether the shot worked. Natural synchronized sound only, no score and no dialogue. Rain and faint street noise at the entrance fade into the lobby hush. Busker's quiet piano grows more audible as Mike approaches the stage and fades as he passes. The elevator chime sounds just before Harper emerges. The lobby ambience recedes as Mike enters the service corridor, ending in sparse fluorescent room tone.

    I realized it wasn’t really taking my idea and optimizing it, but rather optimizing the non-optimized version. This means a lot of clutter was left in the prompt.

    This was when I decided to build instructions (listed at the top). Here is the prompt I received. Notice how much more readable this is (aside from the constant “must strictly follow” instruction).

    A cinematic 30-second single continuous shot in late afternoon on a lively Paris street with narrow sidewalks, Haussmann-style stone façades, café terraces, parked scooters, and market stalls. Mike must strictly follow @image1. Filmed on an ARRI Alexa as a single continuous Steadicam shot, with the operator walking backward in front of Mike and maintaining a natural human-operated feel through subtle footstep drift, gentle lateral corrections, and organic reframing. Mike can move off-center as foreground pedestrians, café furniture, and market stalls create natural depth and occasional brief occlusion, with the framing breathing between medium-wide and medium shots without obvious zooms. Mike walks down the street with AlbumGuy, who must strictly follow @image2, on his left and Lava Lamp Torso, who must strictly follow @image3, on his right. The lava inside her torso visibly moves slowly and continuously. The three are walking together. Mike is relaxed and quietly amused by the people he encounters, looking naturally from side to side as he notices them, with small knowing smiles and understated reactions rather than walking blankly forward. Mike first passes Rory, who must strictly follow @image4, leaning against a brick wall with a hole in it. Rory gestures to the hole and shrugs. He then passes Eleanor, who must strictly follow @image5 and Android who must strictly follow @image6 who are deep in an emotionally moving conversation. Mike passes Devin, who must strictly follow 
    @image7, and his father, who must strictly follow @image8, smiling together at a market stall looking at a fish. Mike passes Busker, who must strictly follow @image9, is playing guitar. At this point, AlbumGuy stops to listen to the Busker while Mike continues walking with Lava Lamp Torso. Mike then passes Frogman, who must strictly follow @image10, who is sipping a coffee with a croissant on a plate in front of him. At this point, Lava Lamp Torso with Frogman and takes a bite of the croissant. Mike continues forward alone with a shrug. Hacker, who must strictly follow @image11, bumps into Mike and reaches into his back pocket to steal his wallet. She then walks away holding the wallet up and looks back toward the camera with a smirk. After the wallet is taken, the camera swings from the front-facing view into a profile view of Mike as Hacker walks away. This profile reframing reveals the adjacent intersecting street and creates room in the composition for the new street to open up beside Mike. Mike is joined on the left side by Concierge, who must strictly follow  @image12 in her blue suit, and on the right side by Harper, who must strictly follow @image13, in her blank tank top and camo cargo shorts, both walking with Mike. Harper playfully bumps into Mike like old friends reunited. As they reach an intersection, Concierge gestures toward the Eiffel Tower visible in the distance down the intersecting street. Each referenced character appears only during their described encounter. Referenced characters must not appear as background pedestrians before their encounter and must not reappear after Mike has passed them, except AlbumGuy and Lava Lamp Torso, who begin with Mike and remain until they stop at their respective encounters. The camera gives each described character or pair a clear readable moment before continuing forward. No music, only ambient Paris street sound.

  • Prompt for Scene Preview Sheets in Image GPT 2

    One interesting way to use GPT Image 2 and Seedance 2.0 to create movies is to generate Scene Previews.

    Here is a prompt to copy-and-paste into a chat window or as instructions in a custom GPT. Be sure to describe your scene where it says: [Paste scene idea here]. You don’t need a reference image of the character but I used one.

    Prompt for Scene Previews in Image GPT 2

    Act as a cinematic visual development art director.
    
    Create one wide landscape visual reference sheet based on the scene below. The result should look like a professional film production-board page or director’s lookbook sheet — not a poster, comic page, generic mood board, scrapbook, or UI dashboard.
    
    SCENE IDEA:
    [Paste scene idea here]
    
    OPTIONAL REFERENCE IMAGES:
    Use any attached images as production references.
    - If an image shows a character, preserve that person’s identity, age range, face shape, hair, body type, wardrobe logic, props, and material details across the sheet.
    - If an image shows a location or environment, incorporate its architecture, geography, weather, lighting, textures, palette, and spatial feeling.
    - If no references are attached, invent original story-specific characters and environments based on the scene.
    - Do not beautify, redesign, or stylize references unless the scene clearly calls for it.
    
    Create a SINGLE visual reference sheet with a clean editorial layout, off-white background, thin black dividers, compact readable labels, and a structured grid.
    
    Include these sections:
    
    TOP BAR: SCENE PREVIEW
    Add a header reading “SCENE PREVIEW”.
    Include:
    - Cut Count: 5
    - Color Palette: infer a restrained cinematic palette from the scene
    - Environment Fingerprint: a short specific description of the setting and atmosphere
    - Visual Rule: one brief sentence that defines the overall visual consistency of the scene
    
    SECTION 1: CHARACTER REFERENCE
    Show the main character, or main characters if needed, in a production-reference format.
    Include:
    - Front view
    - Side view
    - Back view
    - Facial close-up
    - Side/profile close-up
    - Costume/material details
    - Important prop or accessory detail
    - Palette swatches
    - Brief editorial notes
    
    If there are two important characters, split this into Primary Character and Secondary Character.
    
    SECTION 2: ENVIRONMENT / SET DESIGN
    Show the scene environment in a cinematic production-design format.
    Include:
    - One wide hero environment frame
    - One or two supporting set/location frames
    - A top-down floor plan or blocking diagram
    - Numbered camera positions
    - Character movement arrows
    - Entrances, exits, landmarks, and key set pieces labeled
    
    The floor plan does not need architectural precision, but it must communicate clear geography and scene logic.
    
    SECTION 3: STORYBOARD
    Create a 5-cut storyboard strip showing one continuous scene.
    Each shot must maintain continuity of character, wardrobe, lighting, geography, environment, and emotional progression.
    
    For each cut, include:
    - Cut number
    - Lens choice
    - Duration
    - Camera movement or camera style
    - Shot size
    - One concise sentence describing the action and emotional beat
    
    Use professional cinematography language such as:
    35mm anamorphic, 50mm anamorphic, 75mm anamorphic, 100mm macro, static, track, dolly-in, crane-up, rack-focus, handheld, steadicam, push-in, wide, medium, close-up, insert, extreme close-up.
    
    SECTION 4: LIGHTING / MOOD / STYLE NOTES
    Create a lower strip with supporting visual notes.
    Include:
    - 3 to 5 small lighting reference frames
    - Specific lighting labels
    - Mood keywords
    - Lens/style notes
    - Cinematography notes
    - Optional color swatches
    
    STYLE REQUIREMENTS:
    The sheet must feel cinematic, premium, realistic, carefully color-graded, and genuinely useful for production planning. Keep the layout clean and restrained. The imagery should have naturalistic lighting, realistic lens behavior, shallow depth of field where appropriate, controlled motion blur, strong framing, consistent color palette, consistent character identity, and clear environment continuity.
    
    AVOID:
    - Poster composition
    - Comic-book styling
    - Generic concept art collage
    - Watermarks
    - Fake logos
    - Gibberish text
    - Excessive tiny text
    - Repeated identical images
    - Inconsistent faces, costumes, time of day, or geography
    
    Before creating the sheet, silently decide the emotional arc, palette, environment fingerprint, camera progression, lighting progression, and final storyboard beat.
    
    Output one complete cinematic visual reference sheet as a single image.

    I pasted that prompt plus a reference image into ChatGPT and updated the Scene Idea.

    Here is the generated scene preview sheet:

    Then I took both the Scene Preview sheet and the original character reference image and brought them into Dreamina. I think it helps to use both. I imagine using a character reference sheet of the character would be even better than a single image.

    Here is the video:

    Good luck making your movie!

  • Prompts for AI Short Film “Calamity”

    Here are the prompts and images I used to create “Calamity,” a cinematic AI short film made with Seedream 4.5 and Kling 3.0 Omni inside OpenArt Suite. Watch the full workflow — from casting Cleo to building consistent locations, staging a falling piano, and crashing a car through a cafe wall. Below you’ll find the prompts for each scene so you can adapt them for your own projects.

    I share my prompts because I think we all get better when we learn from each other!

    Prompt for Cleo

    A cinematic head and shoulders close-up portrait of a 32-year-old woman with olive skin, hazel eyes, dark hair in a rough graduated bob with deliberately, strong brows, and a calm but intense expression. Natural skin texture with subtle imperfections and faint under-eye fatigue. Shot on a 50mm lens, eye-level close-up framing. Shallow depth of field with softly blurred background. Natural, directional lighting with subtle overhead spill and gentle side contrast. Slightly desaturated color grade. Realistic skin texture with natural imperfections. Soft film grain, high dynamic range, cinematic film still, grounded realism, dramatic but understated.

    Prompt for Raven, the fortune teller

    Cinematic head-and-shoulders portrait of an older woman in her late 60s with a lined, intelligent face and steady presence. Both eyes fully open and facing forward. One eye has a cloudy, milky-white cornea — opaque and desaturated — but the eye is not rolled back, not closed, and not looking upward. The pupil remains centered and natural. The other eye is sharp and observant. Gray hair loosely pinned back. She wears a simple dark wool robe with heavy texture, understated and practical. Calm, unreadable expression with quiet gravity. Natural skin texture with visible age lines. Soft practical interior lighting from a nearby lamp, gentle shadow falloff, muted earthy palette, slightly desaturated. Realistic contemporary photography, 35mm film look, subtle film grain, shallow depth of field.

    Nano Banana Pro Character Sheet prompt

    Character reference sheet of a woman with short dark wavy hair, four full-body views in a row on a clean white background: front view, left profile, right profile, back view. She wears a plain black scoop neck short sleeve shirt, camo cargo pants, and sandals. Relaxed standing pose, consistent identity across all views. Below, three close-up portrait views: front, left profile, right profile. Clean even studio lighting, sharp detail, print-ready technical turnaround sheet.

    Kling 3 Multi-cut prompt for fortune teller scene

    1. Eye-level close-up of the fortune teller one eye clouded milky-white @Raven sitting upright across the table. She leans slightly forward, her brow furrowed with concern. The fortune teller (low, raspy voice, slight Eastern European accent, grave serious tone, slow deliberate pacing): “I am afraid to say I see nothing but calamity.”
    2. Eye-level close-up of @Cleo wearing a sage green tunic, one eyebrow raises skeptically. Cleo (warm clear voice of a young woman with a slight Greek accent moderate pacing): “Calamity?”
    3. Eye-level close-up of the fortune teller sitting upright, she taps a tarot card on the table with one ringed finger and shakes her head slowly. The fortune teller @Raven (low, raspy voice, slight Eastern European accent, somber concerned tone, slow measured pacing): “But that’s not all. Your health. I see something very negative for your health
    4. Close-up of @Cleo as she processes the news, her expression shifting from skepticism to quiet worry. She swallows and looks down at the tarot cards. No dialogue, soft ambient room tone.

    Kling 3 Multi-cut prompt for piano falling scene

    1. Daytime Portland street, overcast light. Medium tracking shot following @Cleo wearing a plain black scoop neck short sleeve shirt, camo cargo pants, and flat sandals. Same outfit in all shots. She walks along a damp sidewalk with shallow puddles. Ahead are two orange traffic pylons near a street post with a bicycle locked to it. A three-story red brick building with storefront windows and an ornate roof cornice runs alongside her.
    2. Low-angle shot that tilts up from the same damp sidewalk past the two orange traffic pylons and the bicycle locked to the street post, up the three-story red brick building façade with storefront windows and an ornate roof cornice. A black grand piano hangs suspended three stories up by a thick rope. The rope snaps and the piano drops downward out of frame.
    3. Medium shot of @Cleo wearing the same plain black scoop neck short sleeve shirt, camo cargo pants, and sandals, walking forward at the same steady pace. Behind her, the piano slams into the pavement with debris scattering. She does not react or turn around.

    Seedream 4.5 prompt in the cafe

    Cinematic interior wide shot of a street-level coffee shop, camera inside near the entrance looking toward a large picture window. @Cleo sits at a small table on the right side of frame, her back to the window and her face turned slightly toward camera. She is wearing a plain black scoop neck short sleeve shirt, camo cargo pants, and flat sandals. On the table: a coffee cup and a croissant. Through the window, a T-intersection is clearly visible: a downhill street runs perpendicular straight toward the café window, and a cross street runs left-to-right parallel to the sidewalk. A red compact car is coming downhill directly toward the café, centered in the window view on a clear collision path. Keep the middle of frame open so the car’s impact path is unobstructed. Realistic coffee shop details: menu board, pastry case or counter in soft focus, other tables and chairs, subtle reflections on the glass. Overcast daylight outside, soft natural interior fill. Muted slightly desaturated colors, 35mm film look, subtle grain, natural lens softness, no HDR, no stylized effects.

    Nano Banana edit prompt to get the car inside the cafe after the crash

    place the car’s hood inside the cafe so the bumper touches the glass pastry case on the left side of the frame and the driver’s side door is directly next to the table with the coffee

    Kling 3 prompt for initial car crash

    @image1 Wide shot from inside facing the large picture window. @Cleo sits at the small table right side of frame, her back to display case on the right, she is wearing a plain black scoop neck short sleeve shirt, camo cargo pants, and flat sandals. The red car approaches with speed, crashing through the window similar to @image2 . @Cleo stays totally unfazed despite the action happening right in front of her. The old woman driving the car reaches out the window and grabs the croissant, taking a bite. @Cleo remains unfazed.

    Kling 3 park bench phone call with her test results

    1. Daytime city park, soft overcast light. @Cleo is wearing a plain black scoop neck short sleeve shirt, camo cargo pants, and sandals sitting alone on a wooden park bench with her iphone next to her, her posture slightly slumped. A few pigeons nearby. Her iPhone vibrates causing the pigeons fly off startled. She looks down at her iphone and sighs, then lifts the phone to her ear.
    2. Closeup profile shot of @Cleo with the phone to her ear. The voice on the other end of the line says, “Hey Cleo, it’s Doctor Patel. We got your test results back. They’re… negative.”
    3. Close up portrait shot of @Cleo. She says (warm clear voice of a young woman with a slight Greek accent moderate pacing and a questioning rising tone): “Negative?” She lowers the phone then repeats peacefully and slowly: “Negative”

    Seedream prompt for Cleo’s apartment

    Wide interior shot of a city apartment, cozy and lived-in with a bohemian, artsy feel. Lots of plants (hanging pothos, potted monstera, small succulents on shelves), mismatched vintage furniture, a worn rug, stacked books, framed prints leaning against the wall, a small record player or speaker, and a soft lamp glow. A couch sits to one side. By the front door: a small table with a key dish. Late afternoon overcast window light mixed with warm practical lamps, soft shadows. Muted earthy palette (warm browns, olive greens, faded textiles), slightly desaturated. Realistic, cinematic, 35mm film look, subtle grain, natural lens softness, no HDR, no stylized effects, no text or watermark.

    Kling 3 prompt for returning home to her cat

    Single shot version

    Interior, bohemian city apartment @image1 . A normal sized cat@image2 is curled up on the couch, relaxed and sleepy. @Cleo enters through the door wearing a plain black scoop neck short sleeve shirt, camo cargo pants, and sandals. She steps in calmly, crosses to the couch and squats beside the cat @image2 , gently pets it. The cat stirs and purrs. @Cleo looks down at the cat with a small, relieved smile and says (warm clear voice of a young woman with a slight Greek accent slow pacing and clear enunciation): “Hello, Calamity.” Lip movements and facial expressions are natural and coherent. The cat @image2 purrs and @Cleo smiles. Camera direction: The camera should smoothly follow @Cleo throughout the shot ending with a close-up on @Cleo and the cat together, ending on her face and the cat’s content expression. Cinematic realism, muted earthy tones, subtle film grain, natural lens softness. No background music.

    Multi-shot version

    1. Interior, bohemian city apartment @apartment. A normal sized cat @Cat is already lying on the couch, relaxed and sleepy. @Cleo enters through the door wearing a plain black scoop neck short sleeve shirt, camo cargo pants, and sandals. She steps in calmly, crosses to the couch and squats beside the cat @Cat, gently pets it. The cat stirs and purrs. Camera direction: The camera should smoothly follow @Cleo
    2. Medium close up @apartment showing @Cat being pet by @Cleo who says (warm clear voice of a young woman with a slight Greek accent slow pacing and clear enunciation): “Hello Calamity.” The cat purrs and she smiles. The camera lingers on the companionship. Cinematic realism, muted earthy tones, subtle film grain, natural lens softness. No background music.

    And of course… Calamity

    cinematic shot of a Bengal cat on a clean plain light-gray seamless background. Shot as studio photography on a 50mm lens at eye-level. Natural directional lighting with subtle overhead spill and gentle side contrast, soft controlled shadows on the floor. Slightly desaturated color grade. Soft film grain, natural lens softness, realistic exposure, cinematic film still look, grounded realism. Print-ready technical turnaround sheet. No props, no environment, no text, no watermark.

    Affiliate links support more free tutorials from AI Video School!

  • How to Create Consistent AI Voices with Text Prompts (No Tools Required)

    Creating consistent voices for your characters helps keep your audience engaged in the story of your AI movie. While the best results often require using multiple tools and additional editing, sometimes you want a method that’s “close enough” and can be done while generating your videos. To be clear, this method does not result in perfect results every time, but it should help you get more consistent voices simply by adding a few things to your prompt.

    To create consistent character voices across multiple shots in AI video, use this format:

    He/She says in the voice of a [AGE] [GENDER], [TIMBRE], [TONE], [PACING]: dialogue

    Examples of AI voice prompts:

    She says in the voice of a middle-aged woman, warm and measured, gentle tone, deliberate pacing: “Thanks for meeting me here.”

    He says in the voice of a weathered middle-aged man, deep and gravelly, matter-of-fact tone, slow pacing: “I knew something was wrong.”

    She says in the voice of a young woman, sharp and clear, dropping to urgent whisper, faster pacing: “No one can know about this.”

    Use My Free Prompt Template!

    Cut and paste the Five Essential Elements for AI Voice Prompts listed below into your favorite AI assistant (ChatGPT, Gemini, Claude, Grok, etc). Then describe the voice or upload an image of the character and ask for a voice prompt. Iterate and refine the prompt. Have fun making your AI film!


    The Five Essential Elements for AI Voice Prompts:

    1. AGE – Approximate age range
      • Examples: young, middle-aged, elderly, teenage, mature
    2. GENDER – Voice register
      • Examples: man, woman, boy, girl
    3. TIMBRE – The physical quality of the voice
      • Examples: deep gentle voice, warm measured voice, sharp clear voice, bright voice, gravelly voice, smooth voice
    4. TONE – The emotional quality or attitude
      • Examples: gentle tone, clinical tone, matter-of-fact tone, urgent whisper, concerned tone, confident tone
    5. PACING – How fast or slow they speak
      • Examples: slow thoughtful pacing, deliberate pacing, moderate pacing, faster pacing, measured pacing

    Pro Tips:

    • Keep AGE, GENDER, and TIMBRE the same across all shots for each character (this is their “voice signature”)
    • Vary TONE and PACING based on emotion (angry = faster, sad = slower, etc.)
    • Be specific – “middle-aged woman, warm and measured” is better than just “woman, nice”
    • Use 2-4 descriptors total after age/gender – more can confuse the AI

    Quick Reference:

    Character signature: [age] [gender], [timbre]
    Current emotion: [tone], [pacing]
    Complete tag: in the voice of a [age] [gender], [timbre], [tone], [pacing]


  • How to Create Consistent Characters with Reference Sheets

    Creating consistent characters is essential for AI filmmaking. Sometimes it’s easy to generate images of a consistent character, but when those images are turned into video, the character starts to look different when they turn or move. What we need to show our AI model is what our character looks like from every angle we plan to show them in.

    I use a technique that allows you to take a single photo of your subject, turn it into a reference sheet, and then use that as an element for generating videos or images with consistent characters. Plus, we can change their wardrobe for different scenes too.

    Reference image

    This is the reference I first tried this with. I wanted a space mechanic with some tattoos and other identifying features. The image itself is dimly lit and it’s not a full body shot. I chose it for those reasons on purpose, for testing purposes.

    Once you have your reference image, generate the character sheet using the prompt at the end of this post. I’m going to use Nano Banana in Google Flow. This should also work if you’re using Nano Banana in an all-in-one tool like Higgsfield, OpenArt, Leonardo, or Freekpik.

    Notice how the tattoo on her neck is the same. In the video, her neck tattoo remains consistent even when she turns around or is off camera then faces camera again.

    Once you have this reference sheet, use it as an element or ingredient in your video generator. This means the generator has to support “references” “ingredients” or “elements,” which are different names for the same thing.

    You can also change the character’s wardrobe with a simple prompt, also at the end of this post.

    Free Consistent Character Prompts

    Here are some prompt templates that I found work well in Nano Banana Pro.

    Consistent Character Prompt with a Reference Image

    Prompt to create a character reference sheet Create a professional character reference sheet based strictly on the uploaded reference image. Use a clean, neutral plain background and present the sheet as a technical model turnaround while matching the exact visual style of the reference (same realism level, rendering approach, texture, color treatment, and overall aesthetic). Arrange the composition into two horizontal rows. Top row: four full-body standing views placed side-by-side in this order: front view, left profile view (facing left), right profile view (facing right), back view. Bottom row: three highly detailed close-up portraits aligned beneath the full-body row in this order: front portrait, left profile portrait (facing left), right profile portrait (facing right). Maintain perfect identity consistency across every panel. Keep the subject in a relaxed A-pose and with consistent scale and alignment between views, accurate anatomy, and clear silhouette; ensure even spacing and clean panel separation, with uniform framing and consistent head height across the full-body lineup and consistent facial scale across the portraits. Lighting should be consistent across all panels (same direction, intensity, and softness), with natural, controlled shadows that preserve detail without dramatic mood shifts. Output a crisp, print-ready reference sheet look, sharp details.

    Consistent Character Prompt with No Reference Image

    Create a professional character reference sheet of [PUT YOUR CHARACTER DESCRIPTION HERE]. Use a clean, neutral plain background and present the sheet as a technical model turnaround in a photographic style. Arrange the composition into two horizontal rows. Top row: four full-body standing views placed side-by-side in this order: front view, left profile view (facing left), right profile view (facing right), back view. Bottom row: three highly detailed close-up portraits aligned beneath the full-body row in this order: front portrait, left profile portrait (facing left), right profile portrait (facing right). Maintain perfect identity consistency across every panel. Keep the subject in a relaxed A-pose and with consistent scale and alignment between views, accurate anatomy, and clear silhouette; ensure even spacing and clean panel separation, with uniform framing and consistent head height across the full-body lineup and consistent facial scale across the portraits. Lighting should be consistent across all panels (same direction, intensity, and softness), with natural, controlled shadows that preserve detail without dramatic mood shifts. Output a crisp, print-ready reference sheet look, sharp details.

    Changing the Wardrobe in Character Sheet in Nano Banana

    Keep the same character sheet layout. Keep their physical characteristics and expression the same. Change the outfit to [outfit description or reference to image].

    Enjoy making your characters more consistent!

  • Free AI Filmmaking Course: Kling, Veo, Nano Banana

    Making a cinematic movie is more than just typing a prompt and letting AI do all the work. It involves combining multiple steps into a cohesive workflow. This free AI mini-course shows you how to create a cinematic short film using Kling and Veo.

    First, let’s watch the movie you’ll learn how to make “Discarded Companion,” created using Kling 01, Veo 3.1, and ElevenLabs.


    Learn the tools: Kling O1, Nano Banana, Google Flow

    Before making “Discarded Companion,” I started comparing Kling O1 to Google Flow. Kling O1 is an all-in-one approach to using elements to create images and video. Google Flow combines Veo and Nano Banana (image generation and editing) into one easy to use tool.

    In this discovery process, I realized the capabilities of the tools opened up the possibilities for even more realistic and engaging storytelling. If you don’t know how to use Kling or Google Flow, this should set you off on the right path to knowing the essentials.


    My in-depth AI video workflow for Discarded Companion

    After making “Discarded Companion,” I showed how I generated every character and scene in the movie. This includes how to edit clips together for longer sequences. The video contains chapters so you can find the section you’re most interested in learning about.

    Start using the AI filmmaking tools you need now!

    Some links on this site are affiliate links. If you purchase through them, I may earn a small commission at no extra cost to you. This helps support AI Video School and allows me to keep creating free tutorials. I only recommend tools I actually use in my filmmaking workflows.