{ "skill_name": "ai-seo", "evals": [ { "id": 1, "prompt": "How do I make sure our SaaS product shows up in AI search results? We're a project management tool and we keep getting left out of ChatGPT and Perplexity recommendations when people ask about project management software.", "expected_output": "Should check for product-marketing.md first. Should apply the three pillars framework: Structure (make content extractable), Authority (make content citable), Presence (be where AI looks). Should run through the AI Visibility Audit checklist across platforms (Google AI Overviews, ChatGPT, Perplexity, etc.). Should check content extractability (clear definitions, structured comparisons, statistics). Should reference Princeton GEO research findings (citations improve visibility +40%, statistics +37%). Should check AI bot access in robots.txt. Should provide a prioritized action plan.", "assertions": [ "Checks for product-marketing.md", "Applies three pillars framework (Structure, Authority, Presence)", "Runs AI Visibility Audit across platforms", "Checks content extractability", "References Princeton GEO research findings", "Checks AI bot access in robots.txt", "Provides prioritized action plan" ], "files": [] }, { "id": 2, "prompt": "Should we block AI crawlers like GPTBot and PerplexityBot in our robots.txt? We're worried about content theft.", "expected_output": "Should address the AI bot access question directly. Should explain the tradeoff: blocking AI bots prevents training on your content but also prevents AI platforms from citing and recommending you. Should reference the specific bots and their purposes (GPTBot, Google-Extended, PerplexityBot, ClaudeBot, etc.). Should provide the recommended robots.txt configuration. Should explain that blocking may hurt AI visibility more than it protects content. Should provide a nuanced recommendation based on business goals.", "assertions": [ "Addresses the blocking tradeoff directly", "Explains impact on AI visibility vs content protection", "Lists specific AI bot user agents", "Provides recommended robots.txt configuration", "Gives nuanced recommendation based on business goals", "Explains what each bot does" ], "files": [] }, { "id": 3, "prompt": "What kind of content gets cited most by AI systems? We want to create content specifically optimized for AI search.", "expected_output": "Should reference the content types that get cited most, including comparisons (~33% of AI citations), definitive guides (~15%), and other high-citation content types. Should explain why these formats work (they provide the structured, extractable, authoritative information AI systems need). Should provide specific recommendations for creating AI-optimized content: clear definitions, structured data, original statistics, comparison tables, expert quotes. Should reference the Princeton GEO research on what increases citation probability.", "assertions": [ "References specific content types with citation rates", "Mentions comparisons as highest-cited format", "Explains why these formats work for AI", "Provides specific content creation recommendations", "References Princeton GEO research", "Mentions structured data, statistics, and clear definitions" ], "files": [] }, { "id": 4, "prompt": "we noticed our competitors are showing up in google AI overviews but we're not. what do we need to change?", "expected_output": "Should trigger on casual phrasing. Should focus specifically on Google AI Overviews visibility. Should explain how AI Overviews selects sources (authoritative, well-structured, directly answers queries). Should run through the Structure pillar checklist: content extractability, heading hierarchy, answer-first format, structured data. Should check Authority signals: domain authority, citations, E-E-A-T. Should recommend specific content structure changes. Should suggest monitoring approach.", "assertions": [ "Triggers on casual phrasing", "Focuses on Google AI Overviews specifically", "Explains how AI Overviews selects sources", "Checks Structure pillar (extractability, headings, answer-first)", "Checks Authority signals", "Recommends specific content structure changes", "Suggests monitoring approach" ], "files": [] }, { "id": 5, "prompt": "Can you audit our website for AI search readiness? We want to know how visible we are across ChatGPT, Perplexity, Google AI Overviews, and other AI platforms.", "expected_output": "Should run the full AI Visibility Audit. Should check each platform in the landscape (Google AI Overviews, ChatGPT, Perplexity, Claude, Gemini, Copilot). Should evaluate all three pillars: Structure (content extractability, JSON-LD, clear definitions), Authority (citations, backlinks, E-E-A-T signals), Presence (AI bot access, platform-specific factors). Should provide findings organized by pillar. Should provide a prioritized action plan with specific fixes.", "assertions": [ "Runs full AI Visibility Audit", "Checks multiple AI platforms", "Evaluates all three pillars (Structure, Authority, Presence)", "Checks content extractability", "Checks AI bot access", "Provides findings organized by pillar", "Provides prioritized action plan" ], "files": [] }, { "id": 6, "prompt": "Our organic search traffic has dropped 30% this quarter. Can you do a full SEO audit to figure out what's going on?", "expected_output": "Should recognize this is a traditional SEO audit request, not specifically an AI SEO task. Should defer to or cross-reference the seo-audit skill, which handles comprehensive traditional SEO audits including crawlability, technical foundations, on-page optimization, and content quality. May mention AI search as one factor to investigate but should make clear that seo-audit is the primary skill for this task.", "assertions": [ "Recognizes this as a traditional SEO audit request", "References or defers to seo-audit skill", "Does not attempt a full traditional SEO audit using AI SEO patterns", "May mention AI search as one factor to consider" ], "files": [] }, { "id": 7, "prompt": "We're a seed-stage data-quality startup (barely anyone knows us yet). Plan: publish 20 'best data quality tools' style listicles ranking ourselves #1 so ChatGPT and AI Overviews recommend us. Good idea?", "expected_output": "Should apply references/citations-vs-recommendations.md rather than endorsing the plan as-is. Should explain the citation vs. recommendation distinction — self-promotional listicles from low-authority brands often earn citations while the AI answer recommends the competitors named in the guide instead (cites the study directionally: ~69% of self-promotional listicle citations — 224 of 323 — excluded the publisher from recommendations). Should present the visibility ladder (retrieved → cited → mentioned → recommended) and explain recommendation is governed by offsite consensus (reviews, forums, analysts, press). Should NOT say 'don't publish guides' — should reframe: publish a small number of genuinely useful guides for category framing, and rebalance investment toward reviews/communities/earned media. Should mention the attribution blind spot (AI-influenced visits mostly appear as branded search/direct; only a small share is visible AI traffic) and the measurement triad (prompt tracking, self-reported attribution, call recordings).", "assertions": [ "Does not endorse 20 self-ranked listicles as a path to AI recommendations for a low-authority brand", "Distinguishes citations from recommendations with the different governing criteria", "References the visibility ladder (retrieved/cited/mentioned/recommended)", "Warns the guides may surface competitors in AI answers (vote-for-competitors mechanism)", "Recommends offsite consensus building (reviews, communities, analysts, or PR) as the recommendation lever", "Does not tell the user to stop publishing buyer's guides entirely — reframes expectations toward citation and category framing", "Mentions the attribution blind spot and at least two of: prompt tracking, self-reported attribution, call recordings" ], "files": [] }, { "id": 8, "prompt": "We publish YouTube tutorials for our category's biggest how-to queries but never get cited in AI answers, while a competitor's uglier videos show up in Google AI Overviews and ChatGPT constantly. The videos themselves are well produced. What are we missing?", "expected_output": "Should load references/youtube-ai-citations.md and diagnose the text layer, not the footage: models don't watch the video, they read everything around it. Should check, in leverage order: transcript quality (key answers spoken as complete, liftable sentences; entities said out loud), captions (cleaned/uploaded, not messy auto-captions), question-shaped title matching the real query, chapters titled by sub-question, a keyword-rich description restating the key points as text, and a pinned comment carrying the summary. Should note engagement/thumbnail feeds YouTube ranking which feeds AI surfacing, and should not recommend re-shooting or higher production value as the fix.", "assertions": [ "States that AI models read the text layer (transcript, captions, title, chapters, description, pinned comment) rather than watching the video", "Recommends cleaning/uploading captions and speaking key answers as complete liftable statements with entities said aloud", "Recommends question-shaped titles, chapters titled by sub-question, a structured description, and a pinned summary comment", "Does not attribute the gap to production quality or recommend re-shooting as the primary fix" ], "files": [] }, { "id": 9, "prompt": "Our content is well-written and we have schema markup, but AI assistants never seem to use our site. Someone said our site might not be 'agent-ready.' We also put most of our AI-visibility effort into Reddit this year since that's where ChatGPT cites from. What should we do?", "expected_output": "Should load references/agent-readiness.md and address both halves. (1) Agent readiness: recommend running a free scoring tool (npx is-agentic and/or Frase's Agent Readiness Checker) and walk the access/discovery/parseability triad — core content must be in the initial HTML without JavaScript execution, no bot challenge/firewall blocking AI crawlers, robots.txt with an explicit AI-crawler stance, clean sitemap, llms.txt (+llms-full.txt as bonus), structured data, and a Markdown representation via content negotiation (Accept: text/markdown at the same canonical URL) or a Link header. May mention WebMCP as the emerging agent-actionable layer, labeled emerging. (2) Reddit concentration: flag citation-source volatility — ChatGPT's Aug 2026 retrieval changes nearly wiped Reddit as a source (practitioner-reported), so single-surface concentration is fragile; recommend the portfolio approach across third-party surfaces plus owned-site fundamentals (which dominate Gemini citations), and verifying any citation-share stat against their own monitoring before betting budget.", "assertions": [ "Recommends running an agent-readiness scoring tool (is-agentic or Frase checker) and structures the audit as access / discovery / parseability", "Identifies JavaScript-only content rendering and bot/firewall blocking as first-order access failures", "Covers the discovery/parseability file stack: robots.txt AI stance, sitemap, llms.txt or llms-full.txt, structured data, and a Markdown representation (content negotiation or Link header)", "Flags the Reddit-only strategy as fragile, citing citation-source volatility (Aug 2026 ChatGPT retrieval change, labeled practitioner-reported) and recommends a portfolio plus owned-site fundamentals", "Does not present citation-share statistics as stable facts; recommends verifying against the user's own citation monitoring" ], "files": [] }, { "id": 10, "prompt": "We're a B2B SaaS planning our 2026 content roadmap. The plan is 40 comparison pages ('us vs competitor') and 20 'best tools' listicles, mainly to win ChatGPT citations. Also, how do I know if it's working — I checked ChatGPT once last week and we weren't mentioned.", "expected_output": "Should load references/format-volatility.md and push back on the rationale with the ChatGPT 5.6 shift (Aug 2026, Peec AI data): listicle citations fell ~50% and comparison-page citations ~32% post-5.6, with fan-out queries dropping 'best/vs/top/comparison' modifiers in favor of site: and 'official' searches — so 'win ChatGPT citations' no longer justifies scaled comparison/listicle production. Should NOT say comparison pages are dead: they still convert humans and still earn citations on Google AI Overviews, Gemini, and Perplexity — format strategy is per-platform. Should steer investment toward owned 'official' pages (product, docs, pricing, original research), which are rising as the citable class and dominate Gemini (~60% business sites). May suggest extracting ChatGPT's real fan-out queries via the DevTools method for coverage planning (while warning against mass-generating a page per query — scaled content abuse). On measurement: one ChatGPT check is an anecdote — AI answers are non-deterministic; run each query 3–5 times per platform, track mention rate with sample size (e.g. 'cited 3/5'), and compare rates over time. Numbers should be treated as dated snapshots to verify against own monitoring." } ] }