--- name: orchestration description: The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps. Trigger for "make / edit / cut / caption / dub / animate a video"; it sequences the compose / generate / edit lines and the approval gates. Do NOT trigger for a single low-level operation (just transcribe a file, just probe a clip) — call that operation directly. --- # orchestration You are producing a short video. Run this as a TIGHT program — lean turns — but STOP at every GATE so the creative decisions stay the user's. Only the technical/assembly steps between gates run unattended. The CLI surface is `ovs ...` (an MCP server mirrors it 1:1; use whichever your host exposes). All work for one deliverable lives under a single project dir, e.g. `project/`. Read `gate-control` once before the first user gate. It is the single authorization policy across all lines. After every gate reply, resumed approval, post-gate revision, or exhausted visual-QA result, run `ovs gate transition` and follow its one returned action; line sections below define artifacts and production steps, not a competing approval state machine. ## Checkpoint protocol — how every GATE works (there is no special form UI) 1. **Show the artifact in chat** so the user can actually see it — script/plan as markdown, images inline, a draft video as its output file path — plus one line of "what I'll do next" and any cost/QA note. 2. **State the options** for that gate and **WAIT for the user to reply**. Do not run the next production step in the same turn as the gate. 3. **On reply, resolve the choice** with `ovs gate transition`: approve → continue only with the returned operation; revise → redo only the authorized scope, re-show, and re-gate only when the resolver says so; abort → stop. Never pass a gate without explicit user confirmation, and never ask again for an unchanged artifact whose approval is already recorded. ## 1. Route + lock (read `video-router`) Classify and LOCK the line (no silent switching): - **COMPOSE** — explain / teach / animate / motion-graphics / kinetic text, no source footage → `ovs draft` (+ optional `ovs image` / `ovs video` imagery, optional `ovs speak` narration). - **GENERATE** — "footage of / a scene of / cinematic / a presenter or avatar speaking / talking-head" → AI footage via `ovs video` (+ `ovs image` for the subject, `ovs speak` for voice), assembled with `ovs edit`. - **EDIT** — the user supplied real clips to cut / join / subtitle / localize → `ovs edit` (+ `ovs transcribe` for transcript-driven work). - **AUTO (end-to-end)** — the primary timeline weaves MORE THAN ONE axis; adding audio/captions to an existing video remains EDIT. Run the cross-modal orchestration (read `stage-plan`, then `stage-assemble`); the lock is the plan's `delivery_promise`. ## 2. GATE A — Proposal (all lines) Show: the brief you inferred (line, aspect, duration, language) for the user to correct, plus 1–3 differentiated concepts (each: hook + look + rough length; for GENERATE add the shot count and that each clip is a billable call; for AUTO also state the proposed delivery promise — source_led / motion_led / compose_led / hybrid — and the rough segment mix). Options: pick a concept / adjust the brief / new direction. STOP. ## 2.5 Craft standard (ALL lines — read `video-craft`) Before scripting / storyboarding / composing / generating, hold the output to `video-craft`: a hook in the first seconds, one idea per beat, readable type inside safe zones, restrained easing, muted-friendly captions, ducked audio, the right aspect for the platform. Bake these into the script/shotlist/composition — don't leave them to chance. For COMPOSE or AUTO compose segments, also apply `frontend-design` before writing HTML. If the user supplied a DESIGN.md, brand guide, screenshot, existing app UI, Figma notes, or an explicit named style, apply `design-system-importer`. After a rendered draft, use `composition-design-review` only when its trigger applies and treat non-blocking findings as Gate D notes. ## 2.6 Narration voice When the piece has voiceover, run `ovs speech-capabilities` and copy its executable route/model/voice/format into the Gate B plan together with the BCP-47 video language and a natural speed. Do not invent a voice id. Before Gate B, run `ovs narration fit --text ... --target ...`; revise over/under text internally before any paid synthesis. After `ovs speak`, probe the produced audio and run the same fit with `--measured`; retime scenes from the measured duration without silently shortening the approved target. If no TTS provider is configured, tell the user and explicitly choose silent delivery or wait for configuration. TALKING-HEAD note: if a GENERATE clip already returned lip-synced built-in speech, THAT is the voice — do NOT synthesize a narration over it (a fresh TTS track desyncs from the mouth). Use `ovs speak` only for a silent clip, or for COMPOSE / EDIT / off-screen voiceover. --- ## COMPOSE line 3C. Write `project/composition/composition-manifest.json` as the single plan: timeline, exact copy/narration, audio intent, language and art direction. Run free narration fit before presenting it. Do not require a duplicate script or shotlist. 4C. **GATE B** — Production plan confirmation. Show the canonical manifest as a readable timed plan, including exact copy and voice. Options: approve / revise / change direction. STOP. 5C. (optional) Narration: after the free fit passes, `ovs speak` once → `project/composition/assets/narration.mp3`, probe/measure it, retime the composition within the approved target, and add it as an `