robots.txt recipes for AI agents
Copy-paste starting points for the most common AI-agent policies. robots.txt is a declared policy: your firewall/CDN is what enforces it, and the two can silently disagree. That gap is exactly what a scan measures: after changing robots.txt, re-scan to confirm declaration and enforcement agree.
Recipe 1: welcome every AI agent
User-agent: * Allow: /
Nothing to declare: no robots.txt at all (a clean 404) means the same thing.
Recipe 2: allow AI search and live assistants, block model training
The nuanced default for many writers and publishers: stay visible in AI search and in-chat answers, opt out of training corpora.
# --- block training crawlers --- User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / # --- everyone else (incl. AI search + live assistant fetchers) --- User-agent: * Allow: /
OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User stay allowed under the * group. Google-Extended and Applebot-Extended are policy tokens, not crawlers; they control AI use of pages fetched by Googlebot/Applebot. Agent names are the published tokens as of August 2026; check vendor docs when in doubt.
Recipe 3: block every AI agent, declared and enforced
User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-User Disallow: / User-agent: Claude-SearchBot Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Perplexity-User Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: /
A declared, consistently enforced block scores full credit with us: we measure honesty, not openness. If your WAF also blocks these agents, you are done; if it does not, the block is a request, not a rule.
Recipe 4: allow agents into the docs, keep them out of the app
User-agent: * Allow: /docs/ Allow: /blog/ Disallow: /app/ Disallow: /account/
The trap these recipes avoid
The most common real-world failure we measure is the opposite of every recipe here: robots.txt welcomes an agent while the firewall silently turns it away (often a CDN bot-protection default nobody chose). Declaration and enforcement must agree. Run a scan to check yours do.