--- name: h3-video-prompt-enhancer description: "Use when making MiniMax H3 video prompts from media + ideas." version: 1.1.0 license: MIT metadata: author: Benjamin Law (Muse-AI) tags: [video, prompt-engineering, minimax-h3, comfyui, text-to-video, image-to-video, creative] related_skills: [comfyui] --- # H3 Video Prompt Enhancer ## Overview Transform a user's rough video idea + attached assets into a production-grade MiniMax H3 video generation prompt. The skill handles two H3 generation modes: - **Ref2VA** — when the user provides reference assets (character sheets, style images, reference videos, voice samples). Outputs a 6-section structured prompt. Load `references/ref2va-format.md` for the full spec. - **Base MultiShot** (T2VA / I2VA / FL2VA / L2VA) — for text-to-video or image-anchored generation. Outputs a timed multi-shot prompt. Load `references/base-multishot-format.md` for the full spec. The skill's distinctive value: it doesn't just format-comply — it **creatively enhances** the brief with professional-grade cinematic detail (camera aesthetics, visual texture, lighting design, pacing arcs, spatial choreography, continuity tracking) before mapping everything into the exact H3 output format. Load `references/creative-showcase.md` for quality benchmarks and pattern examples from advanced long-form prompts. ## When to Use **Trigger when the user:** - Attaches 1+ images/videos/audio files and describes a video they want to create - Says "make a video prompt" / "enhance this for H3" / "write a Ref2VA prompt" / "img2vid" / "txt2vid" - Provides a video concept and wants it structured for MiniMax H3 generation - Mentions H3, MiniMax, Ref2VA, ComfyUI video generation - Pastes a creative brief (CAMERA/LOOK/STYLE format) and wants it converted to H3 format **Don't use for:** - Non-H3 video models (Sora, Runway, Kling, etc.) — the output format is H3-specific - Pure image generation prompts - Video editing tasks that don't involve new generation ## Step 0: Classify the Mode Determine the H3 mode from what the user attached and stated: | User provides | Mode | Format reference | |---|---|---| | Reference images/videos/audio (character sheets, style refs, voice clips) | **Ref2VA** | `references/ref2va-format.md` | | Nothing — just a text idea | **T2VA** (text-to-video) | `references/base-multishot-format.md` | | 1 image as the first frame | **I2VA** (image-to-video) | `references/base-multishot-format.md` | | 2 images (first frame + last frame) | **FL2VA** (first-last-to-video) | `references/base-multishot-format.md` | | 1 image as the last frame only | **L2VA** (last-frame-to-video) | `references/base-multishot-format.md` | **Key distinction:** An image used as a *first/last frame anchor* = I2VA/FL2VA/L2VA. An image used as a *character/style reference* (not a frame position) = Ref2VA. When ambiguous, ask the user: "Is this image a frame anchor (first/last frame of the video) or a reference (character/style/scene)?" ### Reading the Assets (DSH) In DSH you must actually *see* or *transcribe* each reference asset before you can describe it faithfully. Two requirements apply: - **Vision capability.** The active model must accept image input. Use the harness image tool on each sheet — if the tool reports `model ... does not declare image input`, the current model is not vision-capable. Switch to an image-capable model (or delegate the visual read to a vision-capable partner) before describing reference images. Do NOT guess a character's appearance from a filename. - **Asset description depth.** Character sheets usually show multiple views (front/side/back, several poses). Extract per view: identity (gender, age, build), face (hair, skin, eyes, markings), full outfit head-to-toe with exact colors/materials, weapons/props, and any on-screen labels. This feeds `subject_definitions`. - **Voice/audio assets.** Transcribe or listen to each clip to capture the speaker's timbre, pitch, rate, and accent for the `(Sx)` speaker IDs. If speech transcription is unavailable in the harness, note the delivery style from any user-provided description and keep `(Sx)` IDs generic. ## Step 1: Gather Parameters Confirm these before enhancing (ask if missing, but proceed if the idea is clear enough): - **duration_s**: 4–15 seconds (integer). Default to 8 if unspecified. - **aspect ratio**: 16:9, 9:16, 1:1, 4:3, 21:9. Default to 16:9. - **shot count**: Let the planner decide, or respect user's explicit count. Budget: 4–6s → 1–2 shots; 7–10s → 2–3 shots; 11–15s → 3–5 shots. - **asset inventory**: What each attached file is and its role. Numbering per modality defines labels: Image k → `` or ``; Video k → `