generated: '2026-08-11' method: probed source: https://weaveapi.dev/robots.txt http_status: 200 file: weaveapi-content-signals.txt summary: >- WeaveAPI serves Content Signals in robots.txt on its marketing/documentation host. The declaration is the Cloudflare "Managed Content" block — WeaveAPI runs behind Cloudflare and has this feature enabled rather than having hand-authored the directives — but it is genuinely served from the provider's own origin at a stable, machine-readable path, and it is the only preference signal about AI use that WeaveAPI publishes anywhere. content_signals: applies_to: "User-agent: *" directives: - signal: search value: 'yes' means: May collect content to build a search index and return hyperlinks and short excerpts. - signal: ai-train value: 'no' means: May NOT collect content for training or fine-tuning AI models. - signal: use value: reference means: AI systems may consume the content by reference. ai_input_declared: false ai_input_note: >- No ai-input signal is present, so RAG / grounding / real-time retrieval is neither granted nor restricted by Content Signal — the robots.txt preamble states this explicitly as the default. rights_reservation: >- The served preamble asserts that any restriction expressed via Content Signals is an express reservation of rights under Article 4 of EU Directive 2019/790 (Copyright in the DSM). disallowed_agents: - Amazonbot - Applebot-Extended - Bytespider - CCBot - ClaudeBot - CloudflareBrowserRenderingCrawler - Google-Extended - GPTBot - meta-externalagent allowed: - user_agent: "*" rule: "Allow: /" sitemap: https://weaveapi.dev/sitemap.xml irony_note: >- Recorded as an observation, not a judgement: WeaveAPI sells access to OpenAI, Anthropic, DeepSeek, Qwen, Kimi, GLM and MiniMax models, and simultaneously blocks GPTBot, ClaudeBot, CCBot and Google-Extended from its own site and declares ai-train=no over its own content. x-evidence: fetched: '2026-08-11' url: https://weaveapi.dev/robots.txt http_status: 200 content_type: text/plain