--- name: seo-geo description: > Audit and improve content for AI Overviews and answer engines, including citability, entity clarity, crawler access, brand signals, and passage structure. user-invocable: true argument-hint: "[url]" license: MIT metadata: author: AgriciDaniel version: "2.4.2" category: seo --- # AI Search / GEO Optimization (May 2026) ## Primary Source: Google's AI Optimization Guide Google's official position, published under Search Central docs: > "From Google Search's perspective, optimizing for generative AI search is > optimizing for the search experience, and thus **still SEO**." Read `references/google-ai-optimization-guide.md` for the full synthesis, myth-busting list (`llms.txt`, chunking, AI-rephrasing, mention-farming, all rejected by Google as ineffective), and the Who/How/Why test for content quality. Audits should frame GEO findings as **SEO fundamentals applied to AI-search surfaces**, not as a separate optimization discipline. When community recommendations contradict Google's primary source, defer to Google and note the contradiction in the report. ## Key Statistics Third-party figures below were not re-verified on 2026-09-23. Quote them only with their source and date, or leave them out. | Metric | Value | Source | |--------|-------|--------| | AI Overviews reach | 2.5 billion+ monthly active users, reported from Google I/O 2026 keynote coverage; not confirmed on a Google-owned source; 200+ countries | Third-party I/O reporting | | AI Overviews query coverage | ~50% of queries (third-party measurement; varies by country) | Industry data | | AI Mode monthly users | 1B+ (Google, I/O 2026, 2026-05-19) | Google (blog.google) | | AI Mode model | Google upgrades it often (Gemini 3.5 Flash became the default on 2026-05-19, and newer Flash models have shipped since); never tie advice to a model | Google (blog.google) | | AI-referred sessions growth | 527% (Jan-May 2025) | Third-party (attributed to SparkToro; not re-verified) | | ChatGPT weekly active users | 1 billion+ (reported August 2026; 900 million in February 2026) | OpenAI, via press reporting | | Perplexity monthly queries | 500+ million | Perplexity | ## Critical Insight: Brand Mentions > Backlinks **Brand mentions correlate 3x more strongly with AI visibility than backlinks.** (Ahrefs study of 75,000 brands, published 2025-12-12; it follows Ahrefs' May 2025 AI Overviews study) | Signal | Correlation with AI Citations | |--------|------------------------------| | YouTube mentions | ~0.737 (strongest) | | Reddit mentions | High | | Wikipedia presence | High | | LinkedIn presence | Moderate | | Domain Rating (backlinks) | ~0.266 (weak) | **Only 11% of domains** are cited by both ChatGPT and Google AI Overviews for the same query, so platform-specific optimization is essential. --- ## GEO Analysis Criteria (Updated) ### 1. Citability Score (25%) **Self-contained answer blocks** are easy for AI systems to quote. Third-party studies suggest roughly 130-170 words; this is a readability heuristic, not a Google requirement (Google's AI optimization guide says you do not need to chunk content for AI). And **~44% of AI citations come from the first 30% of a page** (SE Ranking study), front-load your most citable, self-contained answer rather than burying it below the fold. **Strong signals:** - Clear, quotable sentences with specific facts/statistics - Self-contained answer blocks (can be extracted without context) - Direct answer in first 40-60 words of section - Claims attributed with specific sources - Definitions following "X is..." or "X refers to..." patterns - Unique data points not found elsewhere **Weak signals:** - Vague, general statements - Opinion without evidence - Buried conclusions - No specific data points ### 2. Structural Readability (20%) **92% of AI Overview citations come from top-10 ranking pages**, but 47% come from pages ranking below position 5, demonstrating different selection logic. **Strong signals:** - Clean H1->H2->H3 heading hierarchy - Question-based headings (matches query patterns) - Short paragraphs (2-4 sentences) - Tables for comparative data - Ordered/unordered lists for step-by-step or multi-item content - FAQ sections with clear Q&A format **Weak signals:** - Wall of text with no structure - Inconsistent heading hierarchy - No lists or tables - Information buried in paragraphs ### 3. Multi-Modal Content (15%) Multi-modal content can support selection in AI answers (third-party claims only; no primary source gives a figure). **Check for:** - Text + relevant images - Video content (embedded or linked) - Infographics and charts - Interactive elements (calculators, tools) - Structured data supporting media ### 4. Authority & Brand Signals (20%) **Strong signals:** - Author byline with credentials - Publication date and last-updated date - **Recency**, content under 3 months old is ~3x more likely to be cited in AI answers; pages left stale 6+ months lose citation eligibility (SE Ranking, 1.3M-citation study). A scheduled refresh program is one of the highest-leverage GEO plays. - Citations to primary sources (studies, official docs, data) - Organization credentials and affiliations - Expert quotes with attribution - Entity presence in Wikipedia, Wikidata - Mentions on Reddit, YouTube, LinkedIn **Weak signals:** - Anonymous authorship - No dates - No sources cited - No brand presence across platforms ### 5. Technical Accessibility (20%) **Many AI crawlers fetch raw HTML without running JavaScript** (for example GPTBot and PerplexityBot in public tests), while Googlebot renders JavaScript and feeds AI Overviews and AI Mode. Server-side rendering keeps content visible to all of them. **Check for:** - Server-side rendering (SSR) vs client-only content - AI crawler access in robots.txt - llms.txt presence (reported for completeness; it carries **no weight** in this score, see `references/llmstxt-evidence.md`) - RSL 1.0 licensing terms --- ## AI Crawler Detection Check `robots.txt` for these AI crawlers: | Crawler | Owner | Purpose | Obeys robots.txt? | |---------|-------|---------|---| | GPTBot | OpenAI | **Model training only** (NOT ChatGPT Search) | yes | | OAI-SearchBot | OpenAI | **ChatGPT Search citability** (the crawler that decides it) | yes | | ChatGPT-User | OpenAI | ChatGPT browsing (user-triggered) | "may not apply" per OpenAI (user-triggered) | | ClaudeBot | Anthropic | **Model training only** (NOT Claude's search features) | yes | | Claude-SearchBot | Anthropic | **Claude/Claude.ai search-result citability** (the crawler that decides it) | yes | | Claude-User | Anthropic | Claude browsing on a user's behalf (user-triggered) | **yes** (Anthropic: all three bots honor robots.txt) | | PerplexityBot | Perplexity | Perplexity AI search (not used to crawl for foundation-model training) | yes | | Perplexity-User | Perplexity | Fetches for a user's question (user-triggered) | generally ignores | | CCBot | Common Crawl | Training data (often blocked) | yes | | Bytespider | ByteDance | TikTok/Douyin AI | yes | | cohere-ai | Cohere | Cohere models | yes | | Google-Extended | Google | **Gemini/Vertex training & grounding, and training of the models behind Search gen-AI features** (NOT Google Search inclusion or ranking) | yes | | Google-CloudVertexBot | Google | Site-owner-requested Vertex AI Agent crawls | yes | | Google-Agent | Google | User-triggered agent fetches (agentic browsing for a user) | **no (user-triggered)** | | Google-GeminiNotebook | Google | Fetches individual user-added source URLs (replaced `Google-NotebookLM`, supported until August 2026) | **no (user-triggered)** | | Google Messages | Google | User-triggered fetch | **no (user-triggered)** | | Applebot-Extended | Apple | **Apple Intelligence / generative-AI training data opt-out only** (NOT Siri, Spotlight, or Safari search; does not itself crawl, it labels content already fetched by Applebot) | yes | Sources: [OpenAI crawlers](https://platform.openai.com/docs/bots), [Google crawlers overview](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers), [Anthropic crawler support article](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), [Apple Applebot-Extended support article](https://support.apple.com/en-us/119829). Anthropic's current crawler support article documents only ClaudeBot, Claude-User, and Claude-SearchBot; it does not list `anthropic-ai`, so the previously-unverified `anthropic-ai` row has been removed rather than kept as a guess. **Recommendation:** Allow OAI-SearchBot, Claude-SearchBot, and PerplexityBot for AI search visibility. GPTBot, ClaudeBot, CCBot, and Applebot-Extended are training-only signals -- allow or block them on licensing preference, not on search-visibility grounds. ### Check the right bot for the claim you are making Two pairs are routinely conflated. **Each claim below may only be supported by its own bot's robots.txt status** -- check them separately and report them separately. | Claim you want to make | Bot to check | Bot that does NOT support this claim | |---|---|---| | "Content is citable in ChatGPT Search" | `OAI-SearchBot` | `GPTBot` | | "Content is available for OpenAI model training" | `GPTBot` | `OAI-SearchBot` | | "Content can be used for Gemini/Vertex training & grounding" | `Google-Extended` | `Googlebot` | | "Content is eligible for Google Search / AI Overviews" | `Googlebot` | `Google-Extended` | | "Content is citable in Claude's search features" | `Claude-SearchBot` | `ClaudeBot` | | "Content is available for Anthropic model training" | `ClaudeBot` | `Claude-SearchBot` | | "Content can be used for Apple Intelligence training" | `Applebot-Extended` | `Applebot` | | "Content is discoverable via Siri, Spotlight, or Safari search" | `Applebot` | `Applebot-Extended` | - **`Google-Extended` governs Gemini and Vertex AI training and grounding use, and training of the models behind Search gen-AI features. It does not affect inclusion in ordinary Google Search, or in AI Overviews and AI Mode, both of which are served from the `Googlebot` index.** Never score `Google-Extended` as a "Google Search readiness" signal, and never cite a blocked `Google-Extended` as evidence that a site is missing from Google Search. - **`OAI-SearchBot` is the crawler that determines ChatGPT Search citability. `GPTBot` is OpenAI's separate training crawler.** Checking `GPTBot` access tells you nothing about whether ChatGPT Search can cite the page. A site that blocks `GPTBot` and allows `OAI-SearchBot` is fully citable in ChatGPT Search. - **`Claude-SearchBot` is the crawler that determines citability in Claude's own search features. `ClaudeBot` is Anthropic's separate training crawler** (per Anthropic's crawler support article). Checking `ClaudeBot` access tells you nothing about Claude search citability, and vice versa; report each separately. - **`Applebot-Extended` is a training-data opt-out signal, not a crawler that fetches pages itself.** Per Apple's support article, disallowing `Applebot-Extended` opts a site out of Apple Intelligence / generative-model training use, but the page remains discoverable through Siri, Spotlight, and Safari as long as `Applebot` itself is allowed. Never cite a blocked `Applebot-Extended` as evidence a site is missing from Apple's search surfaces. Do not use these names interchangeably in report prose. When reporting crawler access, name the specific user-agent that was checked and the specific capability it governs. > **Google's user-triggered fetchers generally ignore robots.txt rules** (Google-Agent, Google-GeminiNotebook, Google Messages); OpenAI says robots.txt "may not apply" to ChatGPT-User, while Anthropic's Claude-User honors it. robots.txt cannot block them, use server-side access controls. Google's canonical crawling/robots reference moved to **developers.google.com/crawling** (migrated 2025-11-20); IP-range files now live at `/crawling/ipranges/` and `googlebot.json` was renamed `common-crawlers.json`. Emerging: **Web Bot Auth** (RFC 9421) lets bots authenticate via a `Signature-Agent` header + key directory (used by Google-Agent); reverse-DNS verification remains the fallback. --- ## llms.txt Standard Read `references/llmstxt-evidence.md` for the primary-source evidence (Mueller, Illyes, SE Ranking 300k-domain study, OtterlyAI server-log audit) on why `/llms.txt` is not currently a citation lever for major AI search systems. claude-seo reports presence but assigns no citation-ranking weight. > **Google now states this explicitly.** Google's AI optimization guide, published > 2026-05-15 (llms.txt guidance clarified 2026-06-15, last updated 2026-07-10), says `llms.txt` and other AI-text files are > not needed for Google Search and do not help or hurt visibility or rankings. > They may still serve non-Google systems. Never recommend `llms.txt` as a Google > ranking or citation lever. Source: > developers.google.com/search/docs/fundamentals/ai-optimization-guide **llms.txt** is a community proposal for giving LLMs a curated map of a site; no major AI provider has confirmed using it. **Location:** `/llms.txt` (root of domain) **Format:** ``` # Title of site > Brief description ## Main sections - [Page title](url): Description - [Another page](url): Description ## Optional: Key facts - Fact 1 - Fact 2 ``` **Check for:** - Presence of `/llms.txt` - Structured content guidance - Key page highlights - Contact/authority information --- ## RSL 1.0 (Really Simple Licensing) New standard (December 2025) for machine-readable AI licensing terms. **Backed by:** Reddit, Yahoo, Medium, Quora, Cloudflare, Akamai, Creative Commons **Check for:** RSL implementation and appropriate licensing terms. --- ## Platform-Specific Optimization | Platform | Key Citation Sources | Optimization Focus | |----------|---------------------|-------------------| | **Google AI Overviews** | Strongly ranking-correlated, cites pages that already rank well | Traditional SEO + passage optimization | | **Google AI Mode** (Gemini models, upgraded often) | Weakly ranking-correlated; broader pool (~9 domains cited/query, Ahrefs) | Distinct surface: freshness, entity authority, citable passages beyond position 5 | | **ChatGPT** | Wikipedia (47.9%), Reddit (11.3%) of top-10 cited sources (Profound, 2025-06-05) | Entity presence, authoritative sources | | **Perplexity** | Reddit (46.7%) of top-10 cited sources (Profound, 2025-06-05), Wikipedia | Community validation, discussions | | **Bing Copilot** | Bing index, authoritative sites | Bing SEO, IndexNow | > **Two Google citation engines, not one.** AI Mode and AI Overviews reach the > same conclusion ~86% of the time but cite the same URLs only **13.7%** of the > time (Ahrefs study, 540K query pairs). Treat them as separate surfaces: ranking > well in classic Search feeds AI Overviews, but AI Mode draws from a broader pool > where freshness and entity authority outweigh raw position. Score both. > > **AI Mode is also a booking surface (2026-08-27).** Flight price tracking > with email alerts (180+ countries and territories), hotel booking through > integrated partners, and fares shown in points or miles now happen inside > AI Mode. Travel and hospitality clients should check partner eligibility; > nothing here is a documented ranking change. > > **UX is now unified, surfaces still distinct.** At Google I/O 2026 (2026-05-19) > Google said follow-up questions now flow from an AI Overview into AI Mode, live > worldwide on desktop and mobile, and began rolling out a new intelligent Search > box where AI Mode is available. The *experience* is one flow, but the two citation engines remain > technically distinct (different models/link sets), keep scoring both. ### Citation surfaces & controls in AI Search (2026) Google added many AI citation/source surfaces across AI Overviews **and** AI Mode (May 2026): - **Preferred Sources**, an eligible domain or subdomain can be selected by a user, making its content more likely to appear in that user's Top Stories and eligible for a preferred badge in AI Mode or AI Overviews. This is a **per-user preference**, not a documented general ranking signal. Publishers may offer Google's interactive button or a deeplink, but should not promise a site-wide ranking lift. Since 2026-09-18 the docs also require the site to be included in Search generative AI features (the Search Console "Search generative AI" control) to show as a preferred source in AI Mode and AI Overviews. Source: developers.google.com/search/docs/appearance/preferred-sources - **"Highly Cited" badges**, earned via original primary reporting that other articles cite. - **Community Perspectives**, elevates Reddit/forum/firsthand content. - Inline links, desktop hover **Link Previews**, and prominent link carousels. **Controlling AI-feature appearance:** there is **no AI-specific opt-out file**, but since 2026-08-31 every site has a Search Console control, "Search generative AI" (include by default, exclude, or inherit), that controls eligibility for AI Overviews, AI Mode and generative AI features in Discover; it is not a ranking signal, it is separate from `Google-Extended`, and it is not a training control. Beyond that, appearance is governed by standard preview/index directives, `nosnippet`, `data-nosnippet`, `max-snippet`, `noindex` (distinct from the third-party AI-crawler robots controls above). Source: developers.google.com/search/docs/appearance/ai-features **Search agents (live, not just WebMCP):** Google's "Information Agents" run in the background to monitor topics, plus agentic booking/calling for select categories (announced at I/O 2026 for a summer US rollout; Information Agents start with AI Pro and Ultra subscribers; confirm current availability before promising it), so agent-friendly-page optimization (real interactive elements, accessibility tree, layout stability) now matters for actions, not only citations. Audit that with `/seo agentic` (the `seo-agentic` sub-skill), which also reads Lighthouse's Agentic Browsing fraction. --- ## Google Update Correlation For AI Overviews or AI Mode visibility changes, check the dated product and core-update entries first: `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run seo_updates.py --kind product --kind core --json`. Treat a stale ledger (`freshness.stale`) as incomplete. ## Output Generate `GEO-ANALYSIS.md` with: 1. **GEO Readiness Score: XX/100** 2. **Platform breakdown** (Google AIO, ChatGPT, Perplexity): give a score only for platforms measured with a tool (for example DataForSEO or SE Ranking); otherwise report qualitative readiness and say it was not measured 3. **AI Crawler Access Status** -- report each crawler separately with the capability it governs. Training access (`GPTBot`, `Google-Extended`, `CCBot`, `ClaudeBot`, `Applebot-Extended`) and search citability (`OAI-SearchBot`, `Googlebot`, `PerplexityBot`, `Claude-SearchBot`, `Applebot`) are distinct findings and must never be merged into one line. 4. **llms.txt Status** (present, missing, recommendations) 5. **Brand Mention Analysis** (presence on Wikipedia, Reddit, YouTube, LinkedIn) 6. **Passage-Level Citability** (self-contained answer blocks identified; ~130-170 words is a heuristic, not a Google rule) 7. **Server-Side Rendering Check** (JavaScript dependency analysis) 8. **Top 5 Highest-Impact Changes** 9. **Schema Recommendations** (for AI discoverability) 10. **Content Reformatting Suggestions** (specific passages to rewrite) --- ## Quick Wins 1. Add "What is [topic]?" definition in first 60 words 2. Create self-contained answer blocks (about 130-170 words is a common heuristic) 3. Add question-based H2/H3 headings 4. Include specific statistics with sources 5. Add publication/update dates 6. Implement Person schema for authors 7. Allow key AI crawlers in robots.txt ## Medium Effort 1. Create `/llms.txt` file (optional: ignored by Google Search; may help other AI crawlers) 2. Add author bio with credentials + Wikipedia/LinkedIn links 3. Ensure server-side rendering for key content 4. Build entity presence on Reddit, YouTube 5. Add comparison tables with data 6. Implement FAQ sections (structured, not schema for commercial sites) ## High Impact 1. Create original research/surveys (unique citability) 2. Build Wikipedia presence for brand/key people 3. Establish YouTube channel with content mentions 4. Implement comprehensive entity linking (sameAs across platforms) 5. Develop unique tools or calculators ## DataForSEO Integration (Optional) If DataForSEO MCP tools are available, use `ai_optimization_chat_gpt_scraper` to check what ChatGPT web search returns for target queries (real GEO visibility check) and `ai_opt_llm_ment_search` with `ai_opt_llm_ment_top_domains` for LLM mention tracking across AI platforms. ## Error Handling | Scenario | Action | |----------|--------| | URL unreachable (DNS failure, connection refused) | Report the error clearly. Do not guess site content. Suggest the user verify the URL and try again. | | AI crawlers blocked by robots.txt | Report exactly which crawlers are blocked and which are allowed. Provide specific robots.txt directives to add for enabling AI search visibility. | | No llms.txt found | Note the absence (optional file; Google Search ignores it) and provide a ready-to-use llms.txt template for non-Google AI crawlers. | | No structured data detected | Report the gap and provide specific schema recommendations (Article, Organization, Person) for improving AI discoverability. | ## FLOW Framework Integration For prompt-guided AI content optimization, use `/seo flow optimize `, FLOW's 21 optimize-stage prompts complement GEO's citability and structure analysis with evidence-led AI prompts.