# YouTube Videos That Get Cited by AI YouTube is one of the most-cited third-party surfaces in AI answers — Google AI Overviews and Gemini cite it heavily, and ChatGPT/Perplexity lift from it for how-to queries. The core insight that changes how you produce for it: **Models don't watch your video. They read everything around it.** The citation is earned by the text layer — title, transcript, captions, chapters, description, and comments — not the footage. A mediocre-looking video with a clean, structured text layer beats a beautiful one that's opaque to a crawler. ## The anatomy Work through these in order of leverage: ### 1. The transcript (the real content) This is what the model actually reads. Optimize the *spoken words*: - **Answer questions in complete, liftable sentences.** "The five steps to create an SOP are…" extracts cleanly; a rambling answer spread across three tangents doesn't. - Script or outline the key answers before recording so each core question gets a clear, structured spoken answer in one place. - Say the important terms out loud — the product name, the category, the entities you want associated. If it's only on a slide, the model may never see it. ### 2. Accurate captions Auto-captions are messy — misheard product names, no punctuation, broken sentences — and messy captions are what the model reads if you don't fix them. Upload cleaned captions (or at minimum correct the auto-generated ones). This is the cheapest fix on the list. ### 3. A question-shaped title Models match the title against the user's prompt. "How to Create SOPs That Scale Your Business" beats a clever title every time. Front-load the question or task; save the branding for the channel. ### 4. Chapters and timestamps Chapters let the model (and viewers) jump to the exact answer. Structure = extractability: each chapter title is another labeled, liftable claim about what the video covers. Match chapter titles to the sub-questions people actually ask. ### 5. A keyword-rich, structured description Restate the video's key points *as text* in the description — a short summary, then a bulleted list of what's covered, then resource links. This reinforces the topic and entities in plain crawlable text and gives the model a second, cleaner copy of the answer. ### 6. A pinned comment with the summary An extra liftable text block: pin a comment with the core answer in numbered steps plus the key links. It's indexed, it's structured, and it survives even when viewers never open the description. ### 7. Thumbnail and engagement Engagement isn't read directly by LLMs, but it drives the watch signals that lift YouTube ranking — and YouTube ranking feeds what AI systems surface and cite. The thumbnail's job is the click; the text layer's job is the citation. ## Publishing checklist - [ ] Title is question- or task-shaped and matches a real query - [ ] Key answers spoken as complete, structured statements - [ ] Captions uploaded or corrected (product names spelled right) - [ ] Chapters added, titled by sub-question - [ ] Description restates the key points in text with a bulleted breakdown - [ ] Pinned comment carries the summary + links - [ ] Important entities (brand, category, product) spoken *and* written ## Related - The same "models read the text layer" logic applies to podcasts: episodes get transcribed and show notes get published, so podcast guesting is earned media that compounds in AI answers — see the `public-relations` skill's podcast guest prep reference. - For producing the videos themselves, see the `video` skill. --- *Anatomy pattern from Ross Simmonds / Foundation Inc. ("The Anatomy of a YouTube Video AI Cites," 2026), distilled and extended with credit.*