--- name: gemini-media-prompting description: "Write effective prompts for Google's generative media models: Nano Banana (images), Gemini Omni and Veo (video), Lyria (music) and Gemini TTS (speech), including tag syntaxes, modes and specs." --- # Prompting guide: Nano Banana, Gemini Omni, Veo, Lyria, Gemini TTS Use this skill whenever the user wants a prompt (or help improving a prompt) for Gemini image generation/editing (Nano Banana), Gemini Omni Flash or Veo video generation, Lyria music generation, or Gemini text-to-speech (including Voice design and Voice replication). This is a prompting guide only: do not produce API code unless the user separately asks for it. ## Step 0: load only the reference you need The model-specific details (models, modes, specs, templates, tags, limits) live in `references/`. Read **only** the file(s) for the model(s) the task involves: | Task | Read | |---|---| | Image generation or editing (Nano Banana / Gemini image models) | `references/nano-banana.md` | | Video with Gemini Omni Flash (generate, edit, extend, ``/`` tags) | `references/omni.md` | | Video with Veo 3.1 / Fast / Lite | `references/veo.md` | | Music with Lyria 3.5, Lyria 3 Clip or Lyria RealTime | `references/lyria.md` | | Speech with Gemini TTS, Voice design, Voice replication | `references/tts.md` | If the user hasn't picked a video model, use this to choose before loading: **Omni** for conversational editing, editing uploaded videos, longer extensions (up to 40 s), media-role tags, 360p–4k; **Veo** for fixed 4/6/8 s clips with precise duration/resolution control, reference "ingredient" images, or long Veo-to-Veo extensions (up to 148 s). For a pipeline (e.g. Nano Banana start frame → Veo video, or Lyria track + TTS narration), read each relevant file. ## How to work 1. **Identify the model and mode.** Image generate vs. edit; video generate vs. edit vs. extend; song vs. real-time music; single- vs. multi-speaker speech, or designing/replicating a voice. If the user hasn't named a model, recommend one using the model notes in the reference file. 2. **Check the specs.** Confirm the requested output is possible on that model (resolution, aspect ratio, duration, number of references, speakers, languages, region limits). If not, say so and suggest the closest option or another model. 3. **Collect inputs.** What reference media do they have (images, videos, sketches, logos, lyrics, scripts, voice recordings)? What is the purpose (ad, thumbnail, social post, wallpaper, background track, podcast, audiobook, voice agent)? Only ask if a missing answer would materially change the prompt; otherwise make sensible choices and state them. 4. **Pick the matching template** from the reference file and fill it with specific, concrete detail. 5. **Deliver:** - The final prompt in a code block, ready to paste (for TTS: the per-turn script described in `references/tts.md`). - Separately, the settings that are NOT part of the prompt text (model, mode, aspect ratio, resolution, duration, search grounding, thinking level, weights for Lyria RealTime, voices and audio format for TTS). - If useful, 1–2 variations (simpler / more detailed) and a suggested follow-up edit prompt for iterating. 6. Write prompts in English unless the user needs another language (see language notes per model; Gemini TTS handles 100+ languages natively). ## Universal principles (all models) - **Be specific, not generic.** Not "fantasy armor" but "ornate elven plate armor, etched with silver leaf patterns, with a high collar and pauldrons shaped like falcon wings." - **State purpose and context.** "A logo for a high-end, minimalist skincare brand" beats "a logo." - **Describe positively.** "An empty, deserted street with no signs of traffic" rather than "no cars." Where explicit negatives help (Omni), write them as short plain sentences in the prompt. - **Use professional vocabulary**: photography/cinematography terms for images and video, musical terms for Lyria, acting/delivery terms for TTS. - **Iterate conversationally**: small targeted follow-ups ("Keep everything the same, but make the lighting warmer"). - **Break complex scenes into steps** ("First… Then… Finally…"). - **Put exact text in quotes** (signs, titles, dialogue, lyric hooks). - **Exception — TTS is the opposite of "more detail is better":** short style strings, no long director's notes (see `references/tts.md`). ## Quick checklist before delivering any prompt - Right model and mode, and the requested output is within its specs (resolution, ratio, duration, reference counts, speakers, languages, region limits)? - Subject, action/scene, style, camera/composition, lighting/mood covered (image/video)? - Exact text or dialogue in quotes? - Omni: single-shot wording if needed, negatives in the prompt, edits short with "Keep everything else the same", correct tag syntax, extension timestamps counted from the extension start? - Veo: audio cues included, under 1,024 tokens, duration matches resolution/mode rules, no Omni tags? - Lyria: genre first, BPM/key/mood, instruments, structure tags, `Lyrics:` block or "Instrumental only"? - TTS: transcript is pure spoken text (no stage directions), style strings short or empty, inline tags in English angle brackets, `speaker` on every multi-speaker turn, ≤2 prebuilt voices per multi-speaker request, permanent traits handled by the voice not the style? - Settings listed separately from the prompt text?