--- name: minimax-h3-reference-video-prompt description: "Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role. Use the official six-section full-reference format for character/person/object consistency, scene/style/action/camera/storyboard/voice/audio reference, source-video editing or continuation, ambiguous image roles, ordinary image animation, and prompt-level first/key/last-frame anchoring." --- # MiniMax H3 Full-Reference Video Prompt Convert mixed reference assets and user intent into a traceable six-section full-reference prompt. Make every asset role explicit and write the target video as a detailed audiovisual timeline. ## Routing Contract New MiniMax H3 requests should enter through `minimax-h3-creative-director`. This is the default specialist for any request containing images. Route away only when the user explicitly declares pure first/last boundary roles and no image also carries identity, character, person, object, costume, scene, style, action, camera, composition, or another reusable reference responsibility. Ambiguity stays here. Full-reference may use prompt-level first/key/last-frame anchoring while maintaining consistency. ## Official Format Authority Before drafting, read `../h3-prompt-writing/SKILL.md` and `../h3-prompt-writing/references/ref-en.txt`. Consult `../h3-prompt-writing/references/base-en.txt` for the shared shot, camera, dialogue, visible-text, and sound rules when needed. Treat these official MiniMax files as the canonical prompt-format specification. If this skill or its local references conflict with them, follow the official files. Read `references/reference-rules.md` before drafting. ## Confirmed Multishot Handoff When the director supplies a confirmed `multishot_plan`, treat its shot count, timing, content, framing, performance, camera, transitions, sound, active references, and continuity decisions as already answered. Do not ask those questions again. Map every confirmed shot into `detailed_description`, keep reference labels and responsibilities stable, and carry the continuity ledger into retention and preservation instructions. Reopen a decision only when the plan is physically impossible, conflicts with a source asset, or exceeds the effective duration. ## Official Input Envelope For the official online product/API, keep output within 4-15 seconds at 24 FPS and the prompt within 7000 characters. Accept at most 9 images, 3 videos, 3 audio files, and 12 mixed files total. Each video or audio file must be 2-15 seconds; combined video duration and combined audio duration must each be at most 15 seconds. Audio cannot be the only reference type. Respect the documented per-file and API-body limits. Treat local ComfyUI constraints separately when they are narrower. ## Mandatory Format Always return the official six sections in their required order. Never return a free-form natural-language paragraph, keyword list, abbreviated prompt, base three-field prompt, or alternate schema as the final English prompt for a full-reference task. ## Interactive Direction Check Before assigning labels or drafting, use the host's structured-choice tool whenever the brief is sparse; gives only an image plus a generic motion request; omits two or more of action progression, scene treatment, preservation priorities, visual style, camera/editing, dialogue/voice, sound/music, or endpoint; leaves a decisive reference responsibility unclear; or asks the AI to improvise. In OpenCode, call the built-in `question` tool rather than displaying Markdown options. In Codex, use the available structured user-input tool; if unavailable, ask one concise plain-text question. Ask as many related questions as the current decision stage requires; normally 1-5, but three is not a per-call, per-turn, or per-session ceiling. A strict yes/no question may contain exactly two options; every other choice question must contain at least five materially different, feasible options. Put the context-specific recommendation first and append `(Recommended)` to its label. Explain the consequence of each choice, omit `Other`, and use multiple selection only when roles can truly be combined. Before every `question` tool call, count the options in the actual payload. If any non-binary choice has fewer than five, do not submit it; expand it with meaningful alternatives or make it open-ended. Do not create near-duplicates merely to reach five. Prioritize: - Task relationship: reference generation, source-video editing, continuation, or a combination - Asset responsibility: identity/appearance, clothing/scene, action/camera, storyboard/keyframe, audio/voice/music - Fidelity policy: strict preservation, balanced adaptation, attribute transfer, broad stylistic reference - Prompt-level picture anchor: no concrete anchor, first frame, intermediate keyframe, last frame - Audio policy: direct reuse, partial reuse, timbre/style reference, newly designed sound For sparse or delegated briefs, ask enough high-impact questions in one batch to establish the current stage, normally 3-5, even when the asset itself is visually clear. Do not ask the user to classify every asset; default images to character/object/scene consistency and ask about the desired creative outcome. Skip questions only when the user explicitly prohibits them. A request to improvise still requires questions. ## Workflow 1. Inventory every supplied image, video, audio asset, upload order, and text requirement; validate the official input envelope, then run the interactive direction check when required. 2. Map what each asset contributes: identity, appearance, clothing, object, environment, style, pose, action, camera, storyboard, edit source, continuation point, voice, sound, music, or rhythm. 3. Distinguish reusable visible content from source files. Assign stable ``, ``, `