---
name: h3-video
description: >
OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product
(MiniMax H3 backend). Use for OpenVideo install/pull/status/run, official 3-field prompts,
T2V/I2V/FL2VA, agent-driven video. Brand is OpenVideo — not a bare ComfyUI workflow.
Triggers: OpenVideo, open-video, H3, generate video, T2V, I2V, FL2VA.
---
# OpenVideo skill · v0.1.0
**Brand: OpenVideo** (always lead with this). MiniMax H3 is the model OpenVideo drives.
**Job:** use **OpenVideo** so an agent produces **good product video**, not a random diffusion
click. Quality comes from **prompt craft + correct mode + validated settings**, not secret samplers.
## 0. Paths (env-first)
Recommended sibling layout (any parent dir name; no machine-absolute paths):
```text
parent/
├── open-video/ # THIS product — CLI, backends, skills ← work here
└── lab/ # ComfyUI + weights (not git)
```
| What | Path / env |
|---|---|
| **Product root (CLI, skills)** | `$OPEN_VIDEO_ROOT` (this checkout) |
| **Install / pull / run** | From product root: `./scripts/install.sh`, `python -m open_video …` |
| **Lab engine + weights** | `$OPEN_VIDEO_LAB` / `$H3_LAB` (sibling `lab/` with ComfyUI + `h3_models/`) |
| **Weights env** | `export OPEN_VIDEO_MODELS=$H3_LAB/h3_models` |
| **ComfyUI URL** | `export OPEN_VIDEO_COMFYUI=http://127.0.0.1:8188` (default) |
**Resolve product root** (never hardcode a machine path):
```bash
# Prefer env, else directory of this skill → repo root
export OPEN_VIDEO_ROOT="${OPEN_VIDEO_ROOT:-$(cd "$(dirname "$0")/../.." 2>/dev/null && pwd)}"
cd "$OPEN_VIDEO_ROOT"
```
Lab harness (same ComfyUI) when you keep a sibling runtime:
```bash
export OPEN_VIDEO_LAB="${OPEN_VIDEO_LAB:-$OPEN_VIDEO_ROOT/../lab}"
export H3_LAB="${H3_LAB:-$OPEN_VIDEO_LAB}"
export OPEN_VIDEO_MODELS="${OPEN_VIDEO_MODELS:-$H3_LAB/h3_models}"
```
Product = `open-video/` checkout. `lab/` is runtime only — never the product root.
---
## 1. Quality hierarchy (what actually moves the needle)
| Rank | Lever | Rule |
|---|---|---|
| **1** | **3-field prompt** | Official structure; concrete visible/audible detail; camera type+amplitude+speed; dialogue in `[lang]…` |
| **2** | **Mode** | T2V / I2V / FL2VA chosen correctly from inputs |
| **3** | **Resolution / steps** | **1344×768**, **20 steps**, `res_multistep` + `simple` (defaults in harness) |
| **4** | **Quant** | INT8 on 5090-class (the only tier `pull` installs); lower VRAM → `recommend-quant` may suggest nf4/w4, which are manual + experimental — never auto-installed |
| **5** | **Duration** | 5–10 s sweet spot (max 15 s / shot); multi-shot via cut times or `open-video` director skill |
| **6** | **Review** | Watch / extract frames; fix prompt; re-run. Dual vision gate only when shipping |
**Anti-patterns (low quality):** bare one-liner prompts; abstract mood words only; wrong mode; NVFP4 on 5090; forcing 2K locally; skipping validation.
---
## 2. Best-quality recipe (local max on RTX 5090)
```text
Diffusion: fl2va_pruned_int8_convrot
Text enc: qwen3vl_32b_int8_convrot
VAE: video fp16 + audio fp32
ComfyUI: --lowvram --use-sage-attention # lowvram optional if VRAM ≥ ~22 GiB free policy
Canvas: 1344×768 (16:9), multiple of 32
Sampler: res_multistep | Scheduler: simple | Steps: 20
Duration: 5–10 s (snaps to 17k+5 frames @ 24 fps)
Audio: native 32 kHz stereo — describe in prompt fields
Prompt: official 3-field only
```
Order-of-magnitude on a 32 GB class NVIDIA GPU: ~10–15 min / 10 s clip @ 1344×768, peak VRAM ~22–30 GB (measure on your machine).
---
## 3. Agent procedure (always this order)
### A. Host ready
```bash
cd "${OPEN_VIDEO_ROOT:?set OPEN_VIDEO_ROOT to this product checkout}"
python -m open_video status # or: open-video status / ps
python -m open_video recommend-quant
```
- Weights incomplete → `python -m open_video pull h3` (or `OPEN_VIDEO_MODELS=… pull`)
- ComfyUI down → start lab server:
```bash
cd "${H3_LAB:-$OPEN_VIDEO_ROOT/../lab}"
curl -sf http://127.0.0.1:8188/system_stats || (
mkdir -p logs && cd ComfyUI && nohup ../venv/bin/python main.py \
--listen 127.0.0.1 --port 8188 --lowvram --use-sage-attention \
> ../logs/comfy_server.log 2>&1 &
)
```
### B. Mode
| Inputs | Mode |
|---|---|
| Text only | **T2V** |
| + 1 image | **I2V** (+ instruction line) |
| + first & last image | **FL2VA** (+ alignment line) |
| Multi-ref identity/style | R2V (needs ref2va weights — not default) |
CLI `--mode` tokens: `t2v` · `i2v` · `flf2v` (FL2VA ↔ `--mode flf2v`). R2V has no CLI token yet.
### C. Craft the 3-field prompt (quality lever #1)
```text
[]
integrated_multimodal_description: [Shot 1]