--- name: higgsfield-seedance-2-5 description: "Seedance 2.5 prompt director — the omni-reference dialect. Routes the four generation modes (t2v / omni_reference / video_edit / video_extension), writes explicit @Image/@Video/@Audio reference roles with exclusions, stages 30-second videos into end-state beats, and covers video editing, forward/backward extension, first-last-frame and multi-keyframe control, storyboard grids, blockout rendering, and seamless transitions. Use whenever the user asks for a Seedance 2.5 prompt, mentions Seedance 2.5 / Dreamina / Jimeng, wants a clip longer than 15s on Seedance, wants to EDIT or EXTEND an existing video rather than generate a new one, or supplies more than a handful of image/video/audio references. For Seedance 2.0 (4K, `mode=fast`, or a genre hint) use higgsfield-seedance instead." user-invocable: true metadata: tags: [higgsfield, seedance, seedance-2.5, dreamina, jimeng, omni-reference, video-edit, video-extension, multi-reference, long-video, keyframes, storyboard, blockout, transitions] version: 1.6.0 updated: 2026-09-26 parent: higgsfield --- # Higgsfield Seedance 2.5 Director Seedance 2.5 is a **different dialect from Seedance 2.0**, not a version bump you can prompt through by habit. 2.0 is a reference-driven shot generator with a 4K lane and a genre hint. 2.5 is an **omni-reference production model**: up to 50 reference materials, 30-second native runtime, and three non-generation modes — it can edit a video you already have, and extend one forward or backward from its boundary frame. The prompt grammar changes with it. Reference roles are declared in prose (`@Image 1 defines …`), audio and text get bracket syntax, long videos are staged with explicit end states, and first/last frames live inside `omni_reference` — either as the platform's `start_image` / `end_image` roles or announced **inside the prompt** — never as a separate mode. > **Model split — read this before writing anything.** 2.5 tops out at **1080p** (no 4K > lane) and has no `genre` hint. Since the 2026-09-26 snapshot it **does** expose > `start_image` / `end_image` roles — but only in `omni_reference` mode. If the job needs > 4K or a genre hint, it is a **Seedance 2.0** job — `../higgsfield-seedance/SKILL.md`. > See § Choosing 2.0 vs 2.5. ## QUICK FACTS *Generated-checked block (scripts/build_index.py verifies anchors). Routing aids — read the linked sections for the rules themselves.* - Four modes, picked **before** writing: `t2v` · `omni_reference` · `video_edit` · `video_extension`; the mode changes what the prompt *is* [→](#the-mode-router) - Higgsfield surface: **480p/720p/1080p** (no 4K), duration **4–30s**, `start_image`/`end_image` **only in `omni_reference`**, `t2v` takes zero references, no genre hint, `extension_mode` required for (and only for) `video_extension` [→](#the-higgsfield-parameter-surface) - `video_edit` **ignores** `duration` and `aspect_ratio` and bills by the source video's length; `video_extension` inherits the source's aspect ratio [→](#the-higgsfield-parameter-surface) - Every reference material gets an explicit role **and** an exclusion — "what to use" plus "what not to use"; never let the model infer the mapping [→](#reference-roles--say-what-to-use-and-what-not-to-use) - Each material also declares a **fidelity grade** — full-preserve / partial-preserve / attribute-transfer (name the target) / loose-guide; beat lines name characters (name + one visible marker), never handles [→](#fidelity--say-how-much-of-each-material-must-survive) - Material budget: 30 images / 10 videos ≤30s total / 10 audio ≤30s total, 50 materials max (Dreamina's figures — the platform enforces only the 30-image and 50-material caps) — a platform `start_image`/`end_image` counts against both; stability ranges are 1–8 subjects (images), 1–5 subjects at 5–10s (video/audio) [→](#material-budget) - Multi-reference is a 5-step workflow — map → group → profile → select-by-scene, one line per subject; `@Images 1 through 4 define four characters` is the canonical failure [→](#multi-reference--the-five-step-workflow) - Long videos are **staged**, not paragraphed: one primary change per stage + an explicit **end state**; timestamps allocate a budget, they are not frame-accurate edit points [→](#long-video--stages-and-end-states) - Staging fixes too many EVENTS; two incompatible JOBS in one generation (physics + performance) is a separate cut — split into two prompts and stitch [→](#split-by-job-not-only-by-length) - Bracket syntax: `()` music · `<>` SFX · `{}` dialogue · `【】` subtitles; non-Chinese dialogue needs a language line before the line; a music suppression never goes inside `()` — `(no music)` there reads as a music cue (house inference, the default) [→](#audio-and-text--bracket-syntax) - First/last frames are `omni_reference` work: the platform `start_image`/`end_image` roles **or** an in-prompt declaration (`@Image 1 is the first frame`) — which holds better is unmeasured; keyframes 3+ are always in the prompt; never merge two anchors into one sentence [→](#first-last-frame-and-multi-keyframe-control) - Editing needs a **sole editing master** + edit scope + Timeline Inheritance; extension needs the **boundary frame aligned before** any new content: `MODE-PLAYBOOKS.md` - Storyboard grids, coarse-vs-fine blockouts, one-click video, seamless transitions: `MODE-PLAYBOOKS.md` - **AI-VFX production pipeline** — model-per-asset-class routing, the size-ref frame, location batching, the `omni_reference` v2v lane (source ≥4s, duration = source), the four-batch rule, the slop catalog: `VFX-PIPELINE.md` - Emotion needs 2–4 **observable** cues, not adjectives; niche camera terms get translated into a visible result [→](#emotional-direction-and-camera-terms) - The real-person formula is 7 slots — and slot 1 is **role, never age**: the age-blind engine rule outranks the source guide's `[Age/Race]` label [→](#the-real-person-character-formula) - Hard limits that must not be over-promised (frame accuracy, locked parameters, pixel-identical transitions) [→](#hard-limits--do-not-over-promise-these) - Dreamina-product features that are **not** on the Higgsfield surface — Ultra Long Video 180s, mark-based editing, Clay Renderer [→](#dreamina-only--what-higgsfield-does-not-expose) --- ## Provenance Two primary sources, labelled throughout, plus the secondary sources below. What each label *means* — and the evidence it requires — is the repo-wide legend in `../shared/provenance.md`. | Label | Source | |---|---| | `[OFFICIAL — Dreamina]` | ByteDance's *Dreamina Seedance 2.5 Prompt Guide* + *User Guide* — the model vendor's own prompt doctrine. Prompt grammar is model-side, so it carries across to Higgsfield's hosting. | | `[OFFICIAL — platform]` | Higgsfield's live `models_explore` catalog, snapshot **2026-09-26** (`../../specs/model-specs.json`). Parameters, enums, and media roles come from here and nowhere else. | | `[DREAMINA-ONLY]` | A Dreamina *product* feature with no Higgsfield parameter behind it. Never quote these as things the user can do here. | | `[EMPIRICAL — sd25-pe]` | `sd25-pe`, a Seedance 2.5 skill file. The repo records only a Discord copy (v0.1.0, noted in the v3.33.0 changelog) and **not who wrote it**, so it is not labelled OFFICIAL. Its mapping-priority claim — material content outranks upload order — agrees with the one measurement here (`MODE-PLAYBOOKS.md` § Panel-to-timestamp mapping — board-first vs board-last, identical order adherence, on Ark). | | Secondary labels | `[OFFICIAL — Higgsfield Seedance 2.5 deck]` · `[DEMO — Higgsfield "AI Love Stories" tutorial]` · `[FIELD — AI-vs-VFX]` (the build in `VFX-PIPELINE.md`) · `[EMPIRICAL — MiniMax H3 skill corpus]` · `[EMPIRICAL — nutllwhy/seedance-tvc-director skill]` — each named where it is used. | Where the two disagree about what is *settable*, the platform snapshot wins — it is what the API actually accepts. --- ## The Mode Router Pick the mode first. The same sentence means different things in different modes, and two of the four modes are not generation at all. | The user wants | Mode | What the prompt is | |---|---|---| | A clip from a description, no materials | `t2v` | A scene brief — the core formula below | | A clip built from images / videos / audio they supply | `omni_reference` | A **role map** plus a scene brief | | To change something inside a video they already have | `video_edit` | An **edit order**: master + scope + preserve list | | More footage before or after a video they already have | `video_extension` | A **boundary contract** plus new content | Two rules that fall out of this: 1. **First/last frames, keyframes, storyboard grids, and blockouts are all `omni_reference`.** 2.5 has no separate first/last-frame mode on this platform. The first and last frame can go in two ways: the platform `start_image` / `end_image` roles (allowed **only** in `omni_reference` — the catalog rejects them in `t2v`, `video_edit` and `video_extension`), or ordinary references whose *role sentence* says they are the first and last frame. Which of the two holds the boundary better is **unmeasured**; keyframes 3+ exist only in the prompt form. `[OFFICIAL — Dreamina: "no need to switch to a separate first/last-frame mode"]` · `[OFFICIAL — platform, snapshot 2026-09-26]` 2. **Editing is not regeneration.** If the user wants the shot rebuilt, that is `omni_reference` with the old clip as a motion reference — not `video_edit`. `video_edit` preserves the master's timeline and changes one scoped thing inside it. > **Video-to-video is not automatically `video_edit`.** The field VFX workflow — swap the > person in this plate, keep every other pixel — runs in **`omni_reference` with the source > attached as a video reference**, because that is the lane where `duration` is settable and > must be **matched to the source** (and where the source therefore has to be **≥ 4 s**, the > `duration` floor). `video_edit` is the lane for a scoped change inside a master whose > timeline must survive untouched. Full routing table + the performance-inheritance clause: > `VFX-PIPELINE.md` § Stage 4. `[FIELD — AI-vs-VFX, 2026-08-08]` --- ## The Higgsfield Parameter Surface `[OFFICIAL — platform, snapshot 2026-09-26]` · verify against `../../specs/model-specs.json` before quoting (HARD RULE 3). | Parameter | Values | Notes | |---|---|---| | `mode` | `t2v` · `omni_reference` · `video_edit` · `video_extension` | default `t2v` | | `duration` | 4–30 s | default 5 — **ignored in `video_edit`** | | `resolution` | `480p` · `720p` · `1080p` | default 720p — **no 4K on 2.5** (1080p added by the 2026-09-26 snapshot; not yet field-rated) | | `generate_audio` | bool | default true | | `bitrate_mode` | `standard` · `high` | default standard | | `extension_mode` | `backward` · `forward` | **required** for `video_extension`, **not allowed** otherwise | | aspect ratio | `auto` · `21:9` · `16:9` · `4:3` · `1:1` · `3:4` · `9:16` | ignored in `video_edit`; follows the source in `video_extension` | | media roles | `start_image` · `end_image` · `image_references` · `video_references` · `audio_references` | `start_image` / `end_image` **only in `omni_reference`** | The catalog also carries **reference-count rules** that the table cannot show `[OFFICIAL — platform, CLI rules 2026-09-26]`: - `t2v` takes **zero** references — no image/video/audio references and no start/end frame. - `omni_reference` needs **at least one** reference (a start or end frame counts). - `start_image` / `end_image` are accepted **only** in `omni_reference`. - image references + start frame + end frame **≤ 30**; all references + start + end **≤ 50**. Three consequences worth stating to the user before they spend credits: - **`video_edit` bills by the source video's duration**, and neither `duration` nor `aspect_ratio` is settable — a 20-second master costs a 20-second render no matter how small the edit. - **`video_extension` inherits the source's aspect ratio**; only the extension's *duration* is yours to set. - **No `genre` parameter.** 2.0's genre hint does not exist here — genre lives in the prompt's visual-style clause instead. Preflight the same way as 2.0: ``` python3 scripts/seedance_lint.py --preflight --model seedance_2_5 "" ``` The linter reads the enums out of `../../specs/model-specs.json`, so an out-of-range duration, a 4K request, or a `video_extension` missing its `extension_mode` is caught before the render. --- ## The Core Prompt Formula `[OFFICIAL — Dreamina]` Combine only the parts the shot needs; omit the rest rather than padding empty slots. ``` performs in . The visuals feature . Use . Audio includes . ``` - **Subject + action** is load-bearing — make it concrete. "The man runs" → "the man accelerates into a sprint while his jacket reacts to the airflow." - **Scene and environment** — location, time, weather, spatial relationships, background state. - **Visual style** — lighting, color, materials, texture, mood. Only descriptors that add information; stacked buzzwords ("cinematic, 8K, masterpiece") sample nothing in particular. The named-substitute discipline in `../higgsfield-seedance/SKILL.md` § Prompt-Craft Laws applies unchanged. - **Camera** — shot size, angle, movement, focus subject, transitions. Motion matches the action; it is not decoration. - **Audio** — dialogue, voice characteristics, ambience, SFX, music, synchronized to the visuals. **Generation parameters are not prompt text.** Resolution, duration, and aspect ratio are set on the generation page or via the API — writing them into the prose does nothing except in the modes that auto-lock them, where they are not settable at all. --- ## Reference Roles — Say What to Use *and* What Not to Use The moment there is more than one material, or a material sitting next to a text description, the prompt must state what each material contributes. `[OFFICIAL — Dreamina]` ``` @Image 1 defines 's . @Video 1 defines . @Audio 1 defines 's . ``` Every material that could leak something unwanted gets an explicit exclusion in the same sentence: ``` @Image 2 defines the workbench and window light. Do not use the people in the image. ``` Rules: - **Mappings live in the prompt.** Text labels drawn inside an image are not a mapping, and the model will not infer which person or prop a material represents. - **Video-only references are motion/pacing references by default**, not identity — say so explicitly when you mean otherwise. - **Several views of one subject must say they are one subject**: "All four images define one folding desk lamp. The output must contain only one lamp throughout." Without it the model duplicates the subject. - **When a reference video already carries the motion accurately, state only which attributes to inherit.** Restating every action fights the reference. A blockout or motion video carries motion and spatial structure — not identity — so the prompt still has to define subjects, scene, action, and visual style. - **Never place a reference handle in a shot where that subject is absent** — the same rule as 2.0's tag discipline; the model forces it into frame. - **Character sheets leak their staging.** A sheet's neutral backdrop and multi-view panel layout are the most common character-material leak — the flat gray studio renders as the actual set. Pair every character-sheet role with its own exclusion: *"Do not take the gray backdrop, the panel borders, or the multi-view layout."* - **Beat lines name characters, never handles.** In action/beat prose, a character appears as name + one visible marker at their first appearance in the beat — *"Mira — silver streak, rust-red jacket — crosses the stall line"* — not as `@Image 2`. The model binds by what it can see in the material, and a handle used as a sentence subject is the classic way one character comes back as two people. `[EMPIRICAL — sd25-pe mapping priority, re-derived 2026-08-09: material content outranks upload order]` **Scope:** this is the 2.5 rule. On 2.0 the house convention is the opposite — the acting paragraph *leads* with the character's tag so the model binds the performance to the right person (`../higgsfield-acting/SKILL.md` § Scene adaptation, `../higgsfield-seedance/SKILL.md` § Tag naming). Handles stay in the role map on 2.5 — the `[Characters]` lines, a staging legend (`@A = the BLUE figure`) — and out of the beat prose (`../shared/house-rulings.md` P2-13). Handle *spelling* follows the surface: Dreamina's guide writes upload-order handles with a space (`@Image 1`); the Higgsfield field build writes the upload-order form without one (`@video1`) and named asset tags (`@size-ref`, `@dragon-v3`). Pick one form per project and never mix them in one prompt. ### Fidelity — say how much of each material must survive `[EMPIRICAL — MiniMax H3 skill corpus, re-derived 2026-08-09; cross-model structure, unmeasured on Seedance]` A role says what job a material does; it still doesn't say how much of the material must reach the pixels. Declare one fidelity grade per material, in the same sentence as its role: - **full-preserve** — the subject appears as-is: face, build, wardrobe, all of it. - **partial-preserve** — the named parts survive, the rest is free: *"the jacket and the scar; hairstyle may change."* - **attribute-transfer** — named traits lift onto a **different, named target**: *"apply this fabric's weave and sheen to Mira's coat."* The target must be named — this is the one case a bare role line cannot express, and the one that goes wrong silently. - **loose-guide** — mood, palette, or energy only; nothing is copied literally. ``` @Image 4 defines the brocade fabric — attribute-transfer onto Mira's coat only. Do not carry the garment's cut, the mannequin, or the studio backdrop. ``` ### Material budget `[OFFICIAL — Dreamina]` Hard limits vs the ranges that actually stay stable: | Type | Hard limit | Stable range | |---|---|---| | Images | 30, each ≤4K | 1–8 distinct subjects | | Videos | 10, ≤30s combined | 1–5 subjects, 5–10s each | | Audio | 10 clips, ≤30s combined | only clips directly relevant | | Video-edit source | 1 video + reference images | source ≤20s, 1–5 reference images | 50 reference materials total. On Higgsfield the platform enforces **two** of these caps and counts a `start_image` / `end_image` against both: image refs + start + end ≤ 30, all materials + start + end ≤ 50 `[OFFICIAL — platform, CLI rules 2026-09-26]`. The 10-video, 10-audio and ≤30 s-combined figures are Dreamina's; no platform rule states them — treat them as the vendor's model limits, unverified on Higgsfield. Above the stable ranges (9–12 subjects in images, 6–10 in audio/video, 6–8 edit reference images) generation still works but stability drops and the shot may need several attempts — budget for it, or split the scene. **More than five subjects needing multiple views → one view per image.** Independent view images beat a single collage of views; the collage is the less stable form. **Spend one view on a strong expression, not four resting faces.** `[OFFICIAL — Higgsfield Seedance 2.5 deck, PART 2]` A set of neutral views teaches the model the face at rest and nothing else, so the first line of dialogue invents a mouth. Generate the views on a neutral light-grey ground (shade and mechanism: `../../templates/ad-asset-prep.md` § Design for win rate) and make **one of them a strong expression** — anger, or a wide smile — so the model learns the character's **facial dynamics and teeth structure**, not only the resting face. The canonical four: front view · back view · facial details at neutral · facial dynamics and teeth under strong emotion. Close the set with the identity line (§ Reference Roles) so all four are read as one person. --- ## Multi-Reference — the Five-Step Workflow `[OFFICIAL — Dreamina]` The goal is **not** to cram every reference into one sentence. It is to define the relationships among characters, props, scenes, actions, and audio so the model picks the right material for the right moment. **Step 1 — name and map each subject individually.** One line per subject: ``` corresponds to @Image 1. Use only the appearance, hairstyle, and clothing. corresponds to @Image 2. Use only the appearance, hairstyle, and clothing. corresponds to @Image 3. Use only the structure, material, and color. references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image. ``` > **The canonical failure:** `@Images 1 through 4 define four characters respectively.` > That sentence does not say which image is which character, and the model will guess. **Step 2 — group by type** once several subjects are in play: `[Characters]` → `[Props]` → `[Scenes]` → `[Motion and Audio]`. Add the non-interchange lock to the character group: "Do not interchange these characters' appearances, clothing, actions, positions, or dialogue." **Step 3 — profile any recurring subject.** A character crossing several scenes, or carrying several materials, gets one consolidated block: ``` [Subject Profile: Conservator] Appearance and clothing: @Image 1. Fixed prop: from @Image 5. Locations: and . Motion references: the case-opening motion from @Video 1. Do not use: other characters' clothing. Do not give this character other equipment. ``` **Step 4 — select references by scene.** Per scene, list only the subset actually used, then the event and its end state: ``` Scene 1 | Inspection in the Conservation Lab Use: , , , and the case-opening motion from @Video 1. Event: opens at the workbench and inspects the sample inside. End state: remains on the inner side of the workbench. stays beside the conservator's right hand. ``` **Step 5 — check ownership.** Props belong to exactly one character ("belongs only to "); character count, clothing, and spatial direction stay constant across scenes. --- ## Long Video — Stages and End States `[OFFICIAL — Dreamina]` 2.5 generates up to 30 seconds natively. Anything with several events gets **staged** — one flat paragraph is where dropped beats come from. Each stage carries exactly **one primary state change** and closes on an **explicit, directly visible end state**. Each new stage restates what carries over. ``` [Generation Goal] Generate a