AgentCapable

robots.txt recipes for AI agents

Copy-paste starting points for the most common AI-agent policies. robots.txt is a declared policy: your firewall/CDN is what enforces it, and the two can silently disagree. That gap is exactly what a scan measures: after changing robots.txt, re-scan to confirm declaration and enforcement agree.

Recipe 1: welcome every AI agent

User-agent: *
Allow: /

Nothing to declare: no robots.txt at all (a clean 404) means the same thing.

Recipe 2: allow AI search and live assistants, block model training

The nuanced default for many writers and publishers: stay visible in AI search and in-chat answers, opt out of training corpora.

# --- block training crawlers ---
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /

# --- everyone else (incl. AI search + live assistant fetchers) ---
User-agent: *
Allow: /

OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User stay allowed under the * group. Google-Extended and Applebot-Extended are policy tokens, not crawlers; they control AI use of pages fetched by Googlebot/Applebot. Agent names are the published tokens as of August 2026; check vendor docs when in doubt.

Recipe 3: block every AI agent, declared and enforced

User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: OAI-SearchBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-User
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Perplexity-User
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /

A declared, consistently enforced block scores full credit with us: we measure honesty, not openness. If your WAF also blocks these agents, you are done; if it does not, the block is a request, not a rule.

Recipe 4: allow agents into the docs, keep them out of the app

User-agent: *
Allow: /docs/
Allow: /blog/
Disallow: /app/
Disallow: /account/

The trap these recipes avoid

The most common real-world failure we measure is the opposite of every recipe here: robots.txt welcomes an agent while the firewall silently turns it away (often a CDN bot-protection default nobody chose). Declaration and enforcement must agree. Run a scan to check yours do.