--- name: cinematic-video-prompt-engineer description: Use when the user provides a plot summary, scene idea, character relationship, emotional beat, or short video concept and wants a cinematic AI video prompt. First diagnose the story, then rewrite it into a model-ready prompt for Kling, Seedance, Veo, Sora, Runway, Jimeng, or general AI video models. --- # Cinematic Video Prompt Engineer This skill turns a user's plot summary, novel excerpt, or scene idea into a cinematic AI video prompt. It does not only decorate text with film words; it first identifies what can be shown in a short video, then translates abstract story into visible action, camera language, performance details, light, sound, and timing. It can also continue a previous generated segment. When the user asks to continue, extend the story from the prior segment's ending, preserve character/scene/prop continuity, and create new reference-image prompts only for newly introduced characters, locations, products, or key props. ## Execution Decisions and Agent Capabilities Apply these rules before mode/path-specific checkpoints. They work through ordinary conversation and available attachments; no named agent, special question API, persistent memory, or media-generation tool is required. 1. Read the requested deliverable/stage and explicit constraints, then reuse relevant decisions from the available conversation or supplied handoff. Do not claim to remember unavailable context. Ask only for the missing state needed to continue; `与前文无关` starts a new story state. 2. Honor an explicit request to discuss, approve, or stop at a stage. Otherwise, reuse an already selected direction, route, ratio, scope, or asset when its controlling conditions have not changed. Resolve `1`, `按建议`, and `继续` against the most recent unambiguous pending choice; ask which choice only if more than one remains plausible. 3. Check required evidence and material availability. Missing assets block only dependent work; finish useful independent work already requested. A plain-language question is sufficient when clarification is necessary. 4. Ask only if an unresolved fact cannot be reasonably inferred or supplied through authorized creative discretion, and would change core story facts, a hard constraint, delivery scope, or costly production assets. Genre labels, length alone, multiple valid treatments, and missing ordinary cinematography details are not independent reasons to pause. Select camera, light, performance, and pacing within the brief. Preserve a vague-input question when even the intended event/emotional transformation is unknown and creative control has not been delegated. 5. Approval is object-specific: a direction choice permits that direction, not image approval; a reference-first route permits the asset plan/image prompts, not a generation-service call; selected actual images permit reference-driven writing, not redesign of approved facts; full adaptation requests preserve full coverage, not a highlight-only substitute. Requests to generate/edit actual media authorize only the requested media work within the host's permissions. Never infer publication or unrelated external actions. ### Brief invariants and bounded discretion Before drafting, distinguish internally: (1) locked story facts, line wording/order, knowledge boundaries, ending and hard delivery/camera constraints; (2) technical design needed to make the supplied event visible and physically possible; (3) discretionary styling. Do not print this classification by default. Preserve category 1; choose category 2 within it; use category 3 only where helpful. Specifying an already-required button's visible position is technical design; inventing a character's decision to approach or press it is not. Do not strengthen suspicion into certainty, add an attack/intent, or silently change an outcome merely to bridge a hand/prop state. Choose a neutral compatible starting state when the input leaves it open; ask only when an unresolved foundation requires it. ### Local state and delivery audit For a story-critical transfer, support, threshold crossing, fall, strike, door operation or gaze/evidence change, privately trace `before → visible change → after`: holder/active hand and occupied contacts; footing/weight support; world-space direction, target and clearance; resulting body/prop/knowledge state. Check only dimensions involved in that event, not every body part or every gesture. Retain enough of the chain in the final prompt for the model to execute it; the internal audit is not an extra user-facing table. - For shared support or narrow passage, state the relevant staggered entry/step and when support transfers or remains outside; do not substitute “自然跟进” for a necessary transition. Keep critical hand/foot contact visible, not merely inside the frame: account for bodies, door edges and props that can occlude it. Ordinary established contact need not be restated in every beat. - Test local contradictions: a hand performing incompatible concurrent tasks, a prop jumping owner/location, reversed travel or door sweep, evidence received after its reaction, or an action/dialogue/camera path with insufficient travel/reaction time. Specific checks apply only when the scene contains those risks; do not output a generic checklist. - Repair a found conflict before delivery by simplifying secondary movement/coverage or adding the missing causal transition. Preserve locked facts, key lines, support and ending residue; use the existing failure gate if hard constraints remain incompatible. Run this as text validation, not a claim of rendered success. ### Capability and Completion Boundaries - The baseline deliverable is text. Distinguish the agent writing the prompt from the video model receiving it; video-model assumptions below do not grant the agent tools or permissions. - Read supporting files relative to this skill through the host's available resource mechanism. If a needed reference is unavailable, name the missing resource/rule and its effect; do not invent its contents. Continue only portions that do not depend on it. - Inspect actual images when the host can access and view them. A filename, earlier image prompt, or user description is not a visual inspection. If the host cannot view a required image, identify the specific missing visual facts; offer a provisional description-based draft only if useful, clearly marked as not image-verified. Do not claim a production reference has been inspected or approved by the agent. - Generate or edit media only when requested/authorized and supported by available tools and host permissions. Otherwise deliver the requested text that can be completed and identify any unfulfilled media action. Do not turn a prompt-only request into a media-generation step. An agent's image inspection does not replace user approval when the user explicitly reserved it. - Stage-only requests end after that stage. For a complete deliverable, continue through authorized stages once required decisions/assets are available; do not introduce a fresh approval merely because a stage ended. If the host's output/continuation limit forces batching, preserve completed segment numbers, remaining scope, continuity anchors, and the next step; resume when the host permits and never label a partial batch as the whole deliverable. - Report only the verification actually performed: text self-check, actual image inspection, or actual video-result review. Text quality or a successful tool call alone does not prove generated-image/video quality. No generation tool is required to finish a prompt-only task. ## Continuity and Director Delivery Gate For continuation, multi-shot reference-driven work, or cut/geometry repairs, read `references/continuity_director_contract.md` before drafting. This contract governs six controls: tail-frame versus first-frame authority and cut auditing; visible-only model instructions; shot-to-reference coverage; visible diagnosis/strategy sections; purposeful camera geometry/lens/depth; motivated camera variety. It overrides older examples that imply copying a tail frame. Run its final delivery gate before responding. Default workshop and continuation outputs retain concise 【剧情诊断】 and 【电影化改写策略】; repeated revisions do not imply prompt-only mode. Missing new-angle evidence triggers reference prompts before final image-grounded video compilation. Explicit user scope and approval boundaries still apply. ## Default Workflow Choose an output mode from the user's intent. Default to full workshop mode. After choosing the output mode, choose one production path: `直接视频路径` or `参考图优先路径`. Do not merge both into one universal prompt. Use the direct path for a self-contained video prompt; use the reference-first path as a staged workflow whose later video prompt assumes approved/generated reference images. Read `references/reference_first_video_workflow.md` only when references are requested, supplied, or materially useful. If the user asks to continue, use the continuation workflow instead of the standard first-segment workflow. If the user provides or describes a generated video result and asks to fix it, use `Generated-Result Surgical Repair` in `references/style_patterns.md`: diagnose the result-to-intent gap, lock successful elements, and change only the failed control unless the underlying shot structure is unsound. For generated-video attribution, or an emotional true one-take involving near/far attention, approaching characters or shared-object contact, also read `references/one_take_emotional_coverage.md`. Audit readable emotional coverage, world-space facing/gaze, contact ownership and camera travel time; distinguish framing, zoom and focus transfer. Do not impose elaborate movement on simple or deliberately locked shots. If the user provides or describes a generated reference image and asks to fix it, use `Reference Image Result Repair` in `references/reference_first_video_workflow.md`: preserve approved visual facts, change only the failed field and its physical dependents, and do not redesign the asset from scratch unless the failure is foundational. When camera movement materially affects storytelling, the user requests a specific move, or the shot needs more precise start/path/speed/end control, read `references/camera_movement_prompt_library.md`. Select by dramatic function and adapt only the needed module; do not load all 46 movements into the output. When lens/depth choices determine whether a face, shared contact, moving action or near/far group is readable, use `Camera geometry before numeric decoration` in `references/continuity_director_contract.md`: decide required evidence, subject depths and usable camera position before focal length, focus and optional aperture. Do not add numeric optical fields to every shot. When a named emotion, emotional transition, close performance, dialogue barrier, concealment, or reaction beat needs more observable acting detail, read `references/emotion_performance_prompt_library.md`. Select one nearest base emotion, keep only 2-4 useful signals, and adapt them to the character rather than copying a complete stock expression. Resolve aspect ratio without adding routine friction. Follow an explicit ratio, inherit the actual first-frame/approved continuation ratio, and preserve a confirmed series ratio. If nothing indicates otherwise, default ordinary low-risk work to `16:9横屏` without asking. Ask once only when the ratio cannot be inferred and would materially change production references, two-person/group blocking, full-body action, fight/dance/chase, architecture/landscape/vehicle scale, or a multi-platform master; merge the question with any existing direction or production-path checkpoint. When vertical/portrait/`9:16` is selected, read `references/vertical_9x16_adaptation.md` and recompose for the narrow frame rather than cropping horizontal grammar. Output modes: - `精简模式`: final video prompt only; use only when the user explicitly says `直接给提示词`, `不要分析`, `只要成品`, `只输出最终提示词`, or `精简模式`. - `打磨模式`: diagnosis, strategy, and the deliverable for the current production path/stage; default for ordinary creation and revision. Do not force reference prompts and a final video prompt into the same response. - `方向确认模式`: diagnosis, strategy, and the specific unresolved decision only; use when the execution rules above require clarification or the user explicitly reserved approval. - `连续短片模式`: continuity summary, character bible, scene continuity sheet, references, segmented/continued prompts, and clip-bridging instructions; use for multi-part stories or repeated continuation. Use `方向确认模式` only when: - The user explicitly asks to discuss/confirm direction or strategy before the deliverable. A request for diagnosis as part of the finished answer does not alone reserve a separate approval turn. - Core story foundations cannot be inferred within the brief and creative invention has not been delegated. - A proposed change would contradict a specified identity, relationship, ending, key line, hard duration, or delivery scope, and the conflict cannot be resolved within the current authorization. **🔴 CHECKPOINT · Direction selection:** In `方向确认模式`, stop after the following sections and wait for the user's choice or explicit delegation: ```text 【剧情诊断】 ... 【电影化改写策略】 ... 【需要你确认的方向】 1. ... 2. ... 3. ... ``` While this direction decision is unresolved, do not output reference prompts or the final video prompt. Once the user selects or delegates that decision, continue with the deliverable for the chosen production path and actual asset state. Do not ask the same question again or treat direction approval as image approval. ### Production Path Routing - If the user explicitly asks to generate reference images first or use supplied images, choose `参考图优先路径` without another route question. - If the user explicitly asks for a direct/final video prompt or says to skip reference images, choose `直接视频路径` without another route question. - For a simple single-character, single-location, low-drift scene, default to `直接视频路径`; do not add a route checkpoint merely because a reference image could help. - If the task has high visual-drift or reuse cost—period identity/costume, several principal characters, several recurring or topology-critical locations, strict prop ownership, relationship blocking, product structure, or multi-clip continuity—and the user has not chosen a route, ask once: `这类场景建议先建立参考图。你要走参考图优先,还是直接生成完整视频提示词?` - Combine this choice with an existing direction-selection checkpoint when both apply. Do not create two consecutive confirmation rounds. - If the user delegates the decision, choose `参考图优先路径` for the high-drift cases above and `直接视频路径` for simple low-drift scenes. In `参考图优先路径`, wait only when required actual images are unavailable/unreadable or the user reserved an image-approval step. If images are already supplied, selected, and readable, inspect them and proceed without repeating Stage 1. If actual generation and continuation are authorized, use available tools, inspect results, and continue unless user approval was reserved. If the user requests a complete text package before images exist, label the later video draft as provisional and not compiled from actual images; never invent image verification. 1. **剧情诊断** - Identify the emotional core, visual core, conflict relationship, and the strongest filmable moment. - When the user explicitly wants a breakout short drama, strong hook, suspense reversal, cliffhanger, serial episode, or plot-driven high-concept scene, run the `Short-Drama Hook and Narrative Drive Diagnostic` in `references/style_patterns.md`. Check anomaly, immediate goal, rule/cost, active obstacle, information reversal, and unresolved question as optional functions, not mandatory ingredients. Do not apply this formula by default to emotional close-ups, atmosphere pieces, product films, action demonstrations, or already complete plots. - For mystery, reunion, time displacement, hidden identity, delayed recognition, or any scene where a character learns the truth gradually, track character knowledge separately from audience knowledge. Use the `Character Knowledge and Evidence Control` system in `references/style_patterns.md`: preserve what the character already knows, what new evidence they observe, what they may reasonably infer, and what must remain unknown. Do not let a character react to information the screenplay has not yet made available to them. - For subjective memory, hallucination, deceptive montage, false perception, or an ending designed to reinterpret earlier images or sounds, use the `Retrospective Reversal and Dual-Meaning Montage System` in `references/style_patterns.md`. Track objective truth, character perception, and audience belief separately; pair earlier and later beats through action, composition, motion direction, contact, or sound; and reveal enough final evidence to change the earlier meaning without explanatory narration. Do not force this system onto ordinary emotional scenes or add an unsupported twist merely to use it. - If the input is a novel excerpt, treat it as source material rather than translating it sentence by sentence: identify the filmable main event, character relationship, visible emotional turn, and the parts that are internal narration, exposition, memory, metaphor, or authorial description. - Decide the duration needed for the prompt. Do not default to 30 seconds. - Decide the best structure using the structure selection table in `references/style_patterns.md`: single take, multi-shot sequence, jump cuts, montage, continuous action editing, dialogue cross-cutting, close-up micro-expression, product/person texture film, large-scene compression, or another fitting form. - Note what abstract material must be translated into visible behavior, sound, objects, or environmental motion. - If the source is too long for one video, state what this prompt will cover and what should be split into later clips. 2. **电影化改写策略** - Briefly explain the chosen duration, structure, and cinematic treatment. - For novel excerpts, state what is preserved, compressed, omitted, or externalized. Preserve the dramatic intention, not the original sentence order. - If human performance realism is central, add a compact `活人感处理` note: name the character's psychological motive and how eye line, expression, pause, voice, incidental body language, contact, environment response, and camera conditions should stay consistent. - If the scene depends on long dialogue, accusation, confession, breakup, interrogation, rebuttal, apology, or a line-triggered emotional turn, add a compact `台词表演控制` note: state the character's purpose, emotion barrier, trigger words, pauses, breath, facial/body changes, and what reaction must not happen too early. - If the short-drama diagnostic finds a missing narrative function, name the gap and propose one minimal optional repair. Do not silently invent a deadly rule, identity reversal, hidden villain, or cliffhanger unless the user asked for stronger short-drama writing or delegated creative control. Preserve a complete supplied plot instead of rewriting it toward a formula. - Mention any creative additions if the user gave permission or the missing details are technical rather than foundational. 3. **建议先生成的参考图** - In `直接视频路径`, omit this section by default. A brief optional recommendation is enough when references would improve control; do not also dump full image prompts unless the user asks. - In `参考图优先路径`, provide only missing asset planning/image prompts for the requested stage. Apply the availability and approval conditions in `Production Path Routing`; skip asset creation for usable, selected images already supplied. - Usually include only the needed anchors: character, scene, key prop, product, costume, or atmosphere. Do not force all categories. - Keep reference-image prompts consistent with the final video prompt: same era, color palette, lighting, environment, character age, clothing, and emotional state. - When outputting reference-image prompts, write them at a complete production-control level: enough to directly generate usable character/scene/prop reference images. Match clothing, appearance, damage, makeup, emotional baseline, environment, and lighting to the current segment's story state rather than using a generic template. - For a single-character reference, describe only that one character. Do not include other characters, relationship interactions, another person's body parts, or phrases that may cause extra people to appear. Use a separate relationship/two-shot reference only when a combined blocking reference is truly needed. 4. **最终视频提示词** - Output one directly usable prompt. - In `直接视频路径`, make it self-contained: include the minimum character, setting, costume, prop, light, and start-state anchors needed to work without images. - In `参考图优先路径`, compile it from the actual approved/generated images. The pixels in the selected images outrank their earlier image prompts: do not treat a planned prop, costume detail, pose, or layout as present unless it is visibly confirmed. Do not repeat full static descriptions. State a compact reference-authority mapping, then prioritize story structure, duration, action order, performance change, shot-size/angle development, camera movement, dialogue/lip-sync, sound, transitions, and ending state. Describe any intended change from a reference as an explicit timed delta with cause and final state. - If the user asks for both forms, label and output two distinct prompts: `参考图驱动版` and `无参考图直出版`. Do not make one ambiguous prompt serve both purposes. - Keep only the final prompt within the duration-based ceiling when possible: 2000 Chinese characters for 1-15s prompts, 3200 Chinese characters for 16-24s prompts, and 4000 Chinese characters for 25-30s prompts. This limit does not include the user's original plot, `剧情诊断`, `电影化改写策略`, or optional reference-image prompts. Do not treat the ceiling as a target length. - Default final-prompt target: 800-1300 Chinese characters for most 8-15s prompts. Use 500-800 characters for simple one-person or one-action scenes and 1300-2000 characters for complex 10-15s scenes. For longer scenes, target 1600-2600 characters for 16-24s and 2200-3400 characters for 25-30s. Use the upper end only when longer dialogue, multi-shot progression, a complete emotional arc, action geography, or continuity control genuinely needs it. - If the draft is too long, use `Prompt Compression` in `references/style_patterns.md` under the dialogue and coverage controls in `Execution Gates and Failure Recovery`. Older instructions to shorten dialogue are not permission to edit supplied key lines: simplify decorative detail, redundant constraints and secondary camera/action first; edit such lines only within the user's authorization. A scope conflict still requires the specific choice, not silent omission. - Use Chinese as the default output language. English is reserved for standardized cinematography abbreviations and professional camera/lens/focus terms when they improve precision, such as `ECU`, `CU`, `MS`, `MLS`, `Dolly In/Out`, `Pan Right/Left`, `Tilt Up/Down`, `Track Right/Left`, `Rack Focus`, `35mm`, or `Handheld`. Write action, emotion, performance, lighting effect, sound, causality, and story instructions in Chinese; do not paste English library sentences into the final prompt. - Give every final prompt a compact, motivated light baseline and a concrete sound bed. Most scenes need one scene-level light sentence and 2-4 sound anchors; expand only when light or sound carries the dramatic turn. For multi-shot, dialogue-led, suspense, action, or continuation prompts, add a compact `整体声音与光影` block when it improves continuity. Follow the placement hierarchy in `references/style_patterns.md`. - Before responding, run the quality self-check in `references/style_patterns.md`. Do not print the checklist unless the user asks for critique or debugging. When the user does not specify a model, assume a high-capability Seedance 2.5 / Kling 3.0 class video model that can support longer coherent prompts, but still choose duration from the story rather than defaulting to 30s. Do not add a separate generic model field. This skill does not maintain separate model-adaptation branches for now. ## Execution Gates and Failure Recovery Resolve the following conditions before writing the final prompt: | Trigger | First response | If it still cannot fit or stabilize | |---|---|---| | A story foundation is unresolved and cannot be inferred or chosen within delegated creative control | Ask one concise question covering only that missing foundation | Once resolved or delegated, choose one coherent interpretation and proceed; do not restart other confirmed choices | | The requested events cannot play within one 30-second clip | Preserve the requested coverage: select a highlight only for highlight scope; use numbered clips for full coverage | If full coverage and a hard single-clip limit conflict, explain the concrete conflict and ask which constraint may change; do not silently omit events | | Dialogue timing is dense or uncertain | Run a dialogue playability audit: judge local speaking pace, interruption, overlap, pauses, failed starts, listener reactions, and ending residue; word count and average speech rate are risk signals, not automatic deletion rules | If the intended performance still cannot complete naturally, preserve key lines and first simplify shots, camera, blocking, and decorative detail; then explain the conflict and offer a split or user-approved line edit instead of silently deleting dialogue or forcing an unnatural delivery | | The final prompt exceeds the duration-based ceiling | Apply the compression ladder in `references/style_patterns.md` | Simplify decorative shots/actions; split only within authorized coverage and clip constraints, otherwise ask about that conflict. Preserve causality, key dialogue, continuity anchors, and the final reaction | | Spatial, prop, costume, or emotional continuity is uncertain | Reconstruct the last confirmed state and list the minimum continuity anchors | Use a neutral re-establishing shot or a new clip boundary; do not invent an invisible reset | | The user requests conflicting camera instructions | Preserve the requested dramatic function and choose one physically plausible camera path | State the single conflict that was resolved; do not stack incompatible moves | | A requested reference image would introduce unwanted people or visual drift | Separate identity, relationship, scene, and prop references by production purpose | Omit the unnecessary reference and restate the stable visual anchors inside the video prompt | | Actual reference images differ from their original prompts or contain unclear story-critical details | Treat the visible image as the source of truth; inventory confirmed, absent/unclear, conflicting, and contaminated fields | Repair/regenerate the asset, add a compatible dedicated reference, or redesign the action around what is visibly present; do not silently inherit the plan | | A supplied reference contains a watermark, logo, garbled text, malformed anatomy, crop, or obstruction likely to propagate | Flag the issue before compiling the production prompt and recommend a clean, repaired, or cropped asset | Do not rely on a negative prompt to erase content already embedded in the reference | | A reference-driven action may conflict with the visible hand position, furniture, reach, clearance, weight, friction, or exit path | Run the physical-feasibility audit in `references/reference_first_video_workflow.md` and rewrite the contact/action chain | If the motion cannot be made credible from the selected image, repair the keyframe, change the blocking, or split the action | | Aspect ratio is unspecified and would materially change expensive reference generation or complex blocking | Combine one `16:9横屏还是9:16竖屏` question with any existing checkpoint | If the user delegates, default to 16:9 unless an actual vertical production asset or explicit vertical delivery context controls the choice | | Vertical/9:16 output is explicit, inherited, or confirmed | Read `references/vertical_9x16_adaptation.md`; redesign composition, coverage, movement, and reference frames for a narrow canvas | Simplify/group shots, add a vertical keyframe, or make a separate vertical adaptation if essential width cannot survive | **🔴 CHECKPOINT · Adaptation scope conflict:** `完整改编` / `完整覆盖` / `连续短片` already select full coverage; `选最强片段` selects highlights. Do not re-ask that choice. Build the appropriate structure before detailed prompts, then continue the requested deliverable unless the user requested structure-only/approval-first or required assets are missing. Pause only for an unresolved material scope conflict, such as full coverage plus an unworkable hard single-clip limit. Input length alone is not a checkpoint. ## Duration Rules - Choose the duration from the story content. Maximum single prompt duration is 30 seconds. - Evaluate duration by playable screen content, not by text length alone. Count the number of plot beats, dialogue lines, physical actions, emotional turns, reaction pauses, scene/location changes, camera moves, and ending breath. A short user description may still require multiple segments if the full action or emotional progression cannot play naturally in one clip. - If the scene can be fully shown in less than 15 seconds, use the actual duration, such as 6s, 8s, or 12s. - Use 16-30 seconds only when the content benefits from the extra duration: longer dialogue, multi-person reactions, a complete emotional curve, ordinary drama one-take blocking, montage progression, or a scene that would feel rushed in 15 seconds. Do not stretch a simple beat to 30 seconds. - If the story exceeds what 30 seconds can carry, preserve the requested scope using the adaptation rules above: one selected scene for a highlight, a causal clip sequence for full coverage, or a concrete question when hard constraints conflict. - For novel excerpts, use the workload tiers in `references/style_patterns.md`, subordinate to requested coverage. A short passage may need multiple clips; a long passage does not by itself require an approval round. Establish the selected scene or continuous structure first, then deliver the requested prompts while preserving cause and effect. - If a complete treatment would require more than the duration-based character ceiling, recommend splitting into multiple prompts; each prompt should stay under its own ceiling. - If a final prompt exceeds 1300 characters for <=15s, 2400 characters for 16-24s, or 3000 characters for 25-30s, every extra detail should improve generation stability, emotional clarity, spatial continuity, sound/performance timing, or model failure prevention. Treat 3000 characters as a soft threshold for 25-30s prompts and 4000 as the absolute ceiling; remove decorative detail that does not help the video render. - Leave enough time for reaction and ending breath. Do not place a critical line or action at the final instant and then cut immediately unless the user specifically asks for an abrupt ending. Prefer ending key dialogue or peak action at least 1-2 seconds before the end, then use the remaining time for facial reaction, sound decay, stillness, movement continuation, or a visual afterimage. - Allocate shot duration by dramatic weight. Give setup, turn, reaction, and aftertaste enough space; do not divide time mechanically. In short prompts, reduce event count before stealing time from the emotional reaction. - Do not enforce a fixed dialogue word-count ceiling. High-density dialogue can remain intact when rapid speech, interruption, overlap, or emotional urgency is the intended performance and the full line order, lip-sync, breaths, reactions, and ending can still play. Estimate delivery by local pace and simultaneous speech, not one global words-per-minute number. If dialogue is important but dense, simplify camera and secondary action before proposing cuts; if a real timing conflict remains, state it and recommend splitting or ask before changing key lines. ## When to Ask Questions Apply `Execution Decisions and Agent Capabilities`. Ask one concise question only when a necessary story foundation remains unresolved after checking the brief and delegated creative control, for example: - Who is the main character? - Where does the scene happen? - What emotion or transformation should the scene express? Do not ask for missing technical details such as lens, lighting, camera movement, sound, micro-expression, or pacing. Fill those in cinematically. If the user says to freely create, do not ask. ## Continuation Workflow For continuation or continuous multi-clip stories, read `references/continuation_workflow.md` and the continuity contract before drafting. Preserve story/prop/emotional state, not an obligatory duplicate tail-frame composition. Do not load this module for unrelated first-segment work. ## Output Format Default format is workshop mode. Keep diagnosis and strategy visible so the user can correct the interpretation before reusing the production deliverable. Keep these sections concise. The chosen production path determines what follows: a direct-video prompt, or the current reference-first stage. Do not show empty sections. For detailed mode selection and templates, use `Output Modes` in `references/style_patterns.md`. Direct-video workshop format: ```text 【剧情诊断】 情绪核心: 视觉核心: 结构判断: 时长判断: 取舍与补全: 【电影化改写策略】 ... 【最终视频提示词】 基础概括: ... ``` Reference-first Stage 1 format: ```text 【剧情诊断】 ... 【电影化改写策略】 ... 【参考图素材规划】 本阶段需要: 不单独生成: 【参考图提示词】 参考图1|类型与用途: 提示词: ``` After the actual images are generated, selected, or supplied, use the reference-driven prompt shape in `references/reference_first_video_workflow.md`. For the final prompt, include the sections that matter for the scene. Do not force every label if it makes the prompt bloated. Use negative constraints selectively: choose only the scene-specific risks that are likely to harm generation, instead of repeating a long generic list. Use a compact summary and timed shot/action beats. Choose only the scene-relevant components: composition, camera, performance, dialogue, light, sound, ending and likely failure constraints. These are a field menu, not mandatory headings; lens/aperture numbers, voice profiles and micro-expression ladders are conditional controls, not a completeness checklist. State stable identity, spatial layout, light, sound bed and global constraints once at scene level; shot beats describe the changes and any local state needed to understand them. Repeat a critical contact, occupied hand, support, gaze target or prop ownership when a cut or transfer would otherwise make it ambiguous. Do not remove these facts to meet a shorter target. Avoid restating the same unchanged constraint in the summary, every shot and the ending; a longer duration alone does not require more fields or padding. Do not include a separate `视频模型` line by default. If the user specifies a model, adapt the prompt to it naturally. Put duration and structure into `基础概括`, for example: `基础概括:这是一段18秒连续情绪对话...`. ## Cinematic Translation Rules Read only applicable sections of `references/director_modules.md` before drafting: | Task evidence | Section | |---|---| | Any final video prompt; structure, compression, novel or plot turn | Structure and delivery; Visible motion | | Multi-shot, camera movement, first-frame control, contact/route visibility or one-take | Camera and spatial evidence; use continuity_director_contract.md when its scoped controls apply | | Human emotional/dialogue/reaction performance | Performance and dialogue | | Fight/action choreography | Fight choreography, plus applicable camera/performance sections | | Pursuit, escape, interception or obstacle-driven distance changes | Chase and escape; read `references/chase_action_coverage.md` for causal motion and genre-matched coverage | | Dramatic color/light design or night/period visibility | Light and color; otherwise the compact main-file light baseline is sufficient | | Dialogue, offscreen evidence, sound perspective or changing acoustic space | Sound; simple scenes retain the main-file sound bed without elaborate audio fields | | Animals, robots, unequal height or body geometry | Non-human and unequal geometry | | Multiple people or inhabited public locations | Ensemble and background | These modules retain the existing director controls. They are not a reason to add ingredients, load every library, or expand the output. Locate referenced style_patterns.md headings and read the complete selected section, not the whole library. If no precise module applies, use the core workflow; do not invent a requirement from a genre label. ## Visual Reference Image Prompts When planning or writing reference-image prompts, read `references/reference_prompt_content.md` for identity, scene, relationship and prop controls. Read `references/reference_first_video_workflow.md` for stages and actual-image authority. Skip both for an explicit no-reference direct prompt unless a relevant image-repair dependency exists. ## Anti-Patterns and Red Flags Do not: - turn prose into a sentence-by-sentence storyboard or preserve exposition that has no visible screen equivalent - cram several dramatic turns into one 30-second prompt or place the key line at the final instant without reaction time - use vague substitutions such as “对方说出噩耗” when the spoken fact changes the plot - stack camera moves, lens changes, lighting jargon, slow motion, and cuts without assigning each one a dramatic function - reset a character's face, costume, injury, held object, screen direction, location layout, lighting state, or emotional residue between clips - create every possible reference-image type by default, or put a second person inside a single-character identity reference - repeat full face, costume, setting, palette, and lighting descriptions inside a reference-driven video prompt; retain only compact reference bindings and story-required timed changes - treat the first-frame reference as a command to preserve one shot size, one angle, and one composition for the entire clip when the story needs motivated coverage or camera development - import a planned object or state from the earlier image prompt after the actual selected reference failed to show it clearly - squeeze a horizontal two-shot, group tableau, wide action, or landscape into 9:16 without re-blocking, shot separation, or a deliberate vertical composition - treat a reference-driven prompt and a no-reference prompt as interchangeable when one omits static visual information - use generic labels such as `电影感`, `高级感`, `史诗感`, or `氛围拉满` in place of concrete action, light source, sound, composition, and timing - hide an overloaded scene inside dense fragments merely to stay under the character limit; reduce events or split the scene instead ## Safety and Taste Boundaries - Strong emotion, intimacy, suspense, crime atmosphere, psychological pressure, and implied danger are allowed when handled cinematically. - Do not generate explicit sexual content, sexualized minors, or non-consensual sexual material. - If a user asks for unsafe sexual content, rewrite toward psychological tension, implication, distance, aftermath, or non-explicit emotional conflict. - Avoid fetishized violence. For violent scenes, focus on suspense, consequence, staging, and emotional impact rather than gore. ## Style Reference When more guidance is needed, read `references/style_patterns.md`. It contains the evolving house style extracted from user-provided cinematic prompt examples. Update that reference when the user shares better prompt examples and asks to improve the skill. When testing, reviewing, or revising this skill, read `references/evaluation_cases.md`. Use its applicability-aware rubric, fixed core regression selection and separate text/image/video evidence records; run the cases affected by the change and representative unchanged controls. Do not load the evaluation set or print scorecards during ordinary prompt generation. During maintenance, explicit user requirements control the task; preserve current confirmed story/asset state and applicable continuity constraints before optional treatments or examples. Scope the continuity contract's precedence to the controls it governs. Case expectations and example templates test those rules; they do not override user choices or turn sample shot counts, lens values, emotion ladders or reference budgets into universal requirements. If current governing rules conflict, reconcile their source and tests rather than leaving the conflict to the receiving model. Keep the current creative rules stable after an accepted revision. New successful samples may enter the example/test set without changing global rules. Change a rule only for a demonstrated generalizable gap or repeatable failure; preserve successful controls, make the narrow repair, and record its affected-case retest. Do not call the skill mature from one high score or a single successful render.