--- name: open-video description: Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future). Use when the user wants to turn a concept, script, or reference image into a finished video or multi-shot film — single shots, image-to-video, first-last-frame interpolation, reference-video/audio styling, or stitched long films beyond the 15s model ceiling. Covers prompt crafting, hard validation, ComfyUI-driven generation, vision judging, refine loop, and ffmpeg stitching. --- # open-video — autonomous director skill > **v0.0.1 default for high-quality H3 clips is [`skill/h3-video`](../h3-video/SKILL.md)** > (Ollama for H3 + harness). Use **this** skill when the user wants multi-shot / long-film > director behavior (plan → judge → stitch) beyond a single strong clip. ## 1. What open-video does open-video is the **autonomous director** layer on top of open video models. It turns a concept into a finished film by running the loop no open engine ships natively: **plan → craft → validate → generate → judge → refine → stitch → deliver** - **Model-agnostic core** (`core/`): planner (coherence bible), crafter, validator, judge-loop, stitcher, selector, ref-pack builder. - **Pluggable backends** (`backends//`): baseline = **MiniMax H3** (#1 open model, Arena parity with closed, native stereo audio). Future: Wan2.2 (physics), LTX-2.3 (speed). - **Engine adapter** (`engines/comfyui/`): drives ComfyUI via its HTTP API. open-video is the brain; ComfyUI is the hands. It is NOT a video engine (ComfyUI is the engine) and NOT a model (H3/Wan/LTX are backends). It is the agent brain ComfyUI lacks: judge→refine + multi-shot stitch + coherence planning **as they land**. **v0.0.1 honesty:** prefer [`skill/h3-video`](../h3-video/SKILL.md) for reliable high-quality single clips. Multi-minute film is the flagship *design*. The judge is REAL when env-wired: set `OPEN_VIDEO_VLM_URL` + `OPEN_VIDEO_VLM_MODEL` (+ `OPEN_VIDEO_VLM_KEY`) to any OpenAI-compatible vision endpoint and every shot is scored + diagnosed automatically; with the env unset it is an honest PASS stub — then manual frame review is mandatory. Single open models still cap ~15s/shot; longer output needs multi-shot orchestration (partial). ## 2. Agentic procedure — run these steps in order **Step 1 — Understand the request.** Classify the job: single shot vs multi-shot film; target duration; did the user supply reference image(s) / video / audio; desired aspect; quality bar. If the request is ambiguous *and* the target is a film >30s, ask ONE focused question (subject + mood + length). Do not generate a long film on guesswork. **Step 2 — Pick the mode.** Mode is auto-derived from inputs (see `backends/h3/backend.py` and `scripts/validate_prompt.py` `detect_mode`): - **T2V / T2VA** — text only, no image. Native audio + video. - **I2V / I2VA** — one image (first frame). Instruction line: `"For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced."` - **FL2VA** — two images (first + last frame). Continuous interpolation. **This is the multi-shot chain mode**: previous shot's last frame → next shot's first frame for continuous handoff. - **L2VA** — one image referenced at the **final** timestamp. - **R2V / Ref2VA** — reference video/audio for identity / style / motion / voice. Tag refs ``, `