--- name: stage-generate description: AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO. Generate per shot with `ovs video`, preserve signed reference/settings, then assemble. For recurring characters also read stage-consistency. --- # stage-generate How to produce AI-generated footage and execute a **bounded semantic video edit** already planned by EDIT/AUTO. Designed HTML remains composition work; deterministic cutting remains stage-edit work. The generation tools are `ovs image`, `ovs video`, and `ovs speak`, followed by `ovs edit` assembly. Every billable call belongs to a signed generate segment. ## Pattern A — talking-head / spokesperson 1. **Character still:** generate one image of the presenter / avatar with the intended look. **Keep this reference image** and reuse it for every shot of the same character. 2. **Bring it to life:** generate a video *from* that image (image-to-video). When the provider returns speech + **built-in audio**, that audio is the deliverable voice — it is **lip-synced to the mouth in the clip** — so keep it and do NOT synthesize a separate narration. Only when the clip comes back **silent** do you synthesize the narration (`ovs speak`) and add it as the audio track. Synthesizing a fresh TTS track over a clip that already speaks is the #1 talking-head defect: the new audio has different wording/timing/length, so the voice no longer matches the lips. 3. **Polish:** add captions / a lower-third / a hook by authoring a small composition (stage-compose) and overlaying it onto the clip — **visual-only**. Preserve the clip's own (lip-synced) audio through assembly; a captions composition must not carry a narration `