--- name: h3-video description: > OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend). Use for OpenVideo install/pull/status/run, official 3-field prompts, T2V/I2V/FL2VA, agent-driven video. Brand is OpenVideo — not a bare ComfyUI workflow. Triggers: OpenVideo, open-video, H3, generate video, T2V, I2V, FL2VA. --- # OpenVideo skill · v0.1.0 **Brand: OpenVideo** (always lead with this). MiniMax H3 is the model OpenVideo drives. **Job:** use **OpenVideo** so an agent produces **good product video**, not a random diffusion click. Quality comes from **prompt craft + correct mode + validated settings**, not secret samplers. ## 0. Paths (env-first) Recommended sibling layout (any parent dir name; no machine-absolute paths): ```text parent/ ├── open-video/ # THIS product — CLI, backends, skills ← work here └── lab/ # ComfyUI + weights (not git) ``` | What | Path / env | |---|---| | **Product root (CLI, skills)** | `$OPEN_VIDEO_ROOT` (this checkout) | | **Install / pull / run** | From product root: `./scripts/install.sh`, `python -m open_video …` | | **Lab engine + weights** | `$OPEN_VIDEO_LAB` / `$H3_LAB` (sibling `lab/` with ComfyUI + `h3_models/`) | | **Weights env** | `export OPEN_VIDEO_MODELS=$H3_LAB/h3_models` | | **ComfyUI URL** | `export OPEN_VIDEO_COMFYUI=http://127.0.0.1:8188` (default) | **Resolve product root** (never hardcode a machine path): ```bash # Prefer env, else directory of this skill → repo root export OPEN_VIDEO_ROOT="${OPEN_VIDEO_ROOT:-$(cd "$(dirname "$0")/../.." 2>/dev/null && pwd)}" cd "$OPEN_VIDEO_ROOT" ``` Lab harness (same ComfyUI) when you keep a sibling runtime: ```bash export OPEN_VIDEO_LAB="${OPEN_VIDEO_LAB:-$OPEN_VIDEO_ROOT/../lab}" export H3_LAB="${H3_LAB:-$OPEN_VIDEO_LAB}" export OPEN_VIDEO_MODELS="${OPEN_VIDEO_MODELS:-$H3_LAB/h3_models}" ``` Product = `open-video/` checkout. `lab/` is runtime only — never the product root. --- ## 1. Quality hierarchy (what actually moves the needle) | Rank | Lever | Rule | |---|---|---| | **1** | **3-field prompt** | Official structure; concrete visible/audible detail; camera type+amplitude+speed; dialogue in `[lang]…` | | **2** | **Mode** | T2V / I2V / FL2VA chosen correctly from inputs | | **3** | **Resolution / steps** | **1344×768**, **20 steps**, `res_multistep` + `simple` (defaults in harness) | | **4** | **Quant** | INT8 on 5090-class (the only tier `pull` installs); lower VRAM → `recommend-quant` may suggest nf4/w4, which are manual + experimental — never auto-installed | | **5** | **Duration** | 5–10 s sweet spot (max 15 s / shot); multi-shot via cut times or `open-video` director skill | | **6** | **Review** | Watch / extract frames; fix prompt; re-run. Dual vision gate only when shipping | **Anti-patterns (low quality):** bare one-liner prompts; abstract mood words only; wrong mode; NVFP4 on 5090; forcing 2K locally; skipping validation. --- ## 2. Best-quality recipe (local max on RTX 5090) ```text Diffusion: fl2va_pruned_int8_convrot Text enc: qwen3vl_32b_int8_convrot VAE: video fp16 + audio fp32 ComfyUI: --lowvram --use-sage-attention # lowvram optional if VRAM ≥ ~22 GiB free policy Canvas: 1344×768 (16:9), multiple of 32 Sampler: res_multistep | Scheduler: simple | Steps: 20 Duration: 5–10 s (snaps to 17k+5 frames @ 24 fps) Audio: native 32 kHz stereo — describe in prompt fields Prompt: official 3-field only ``` Order-of-magnitude on a 32 GB class NVIDIA GPU: ~10–15 min / 10 s clip @ 1344×768, peak VRAM ~22–30 GB (measure on your machine). --- ## 3. Agent procedure (always this order) ### A. Host ready ```bash cd "${OPEN_VIDEO_ROOT:?set OPEN_VIDEO_ROOT to this product checkout}" python -m open_video status # or: open-video status / ps python -m open_video recommend-quant ``` - Weights incomplete → `python -m open_video pull h3` (or `OPEN_VIDEO_MODELS=… pull`) - ComfyUI down → start lab server: ```bash cd "${H3_LAB:-$OPEN_VIDEO_ROOT/../lab}" curl -sf http://127.0.0.1:8188/system_stats || ( mkdir -p logs && cd ComfyUI && nohup ../venv/bin/python main.py \ --listen 127.0.0.1 --port 8188 --lowvram --use-sage-attention \ > ../logs/comfy_server.log 2>&1 & ) ``` ### B. Mode | Inputs | Mode | |---|---| | Text only | **T2V** | | + 1 image | **I2V** (+ instruction line) | | + first & last image | **FL2VA** (+ alignment line) | | Multi-ref identity/style | R2V (needs ref2va weights — not default) | CLI `--mode` tokens: `t2v` · `i2v` · `flf2v` (FL2VA ↔ `--mode flf2v`). R2V has no CLI token yet. ### C. Craft the 3-field prompt (quality lever #1) ```text [] integrated_multimodal_description: [Shot 1]