Choose which crawlers to keep out and which paths no one should crawl. The generator writes the file, then checks it crawler by crawler with Google’s matching rules — so blocking AI training does not quietly take you out of AI search.
Checked, not assumed: every crawler run against the file above with the tester’s engine.
Every large AI company now runs more than one crawler, and they do different jobs. The training crawlers collect content for future models. The search crawlers put your pages into answers that link back. Blocking the second kind to stop the first costs you visibility and protects nothing.
| Company | Training — safe to block | Search and answers — blocking costs visibility |
|---|---|---|
| Google-Extended (a token, not a crawler) | Googlebot, which also covers AI Overviews and AI Mode | |
| OpenAI | GPTBot | OAI-SearchBot (ChatGPT search) |
| Anthropic | ClaudeBot | Claude-SearchBot, Claude-User |
| Apple | Applebot-Extended (a token, not a crawler) | Applebot (Spotlight, Siri, Safari) |
| Perplexity | — | PerplexityBot, which Perplexity says is not used for training |
| Meta | Meta-ExternalAgent | Meta-WebIndexer (Meta AI search) |
| Amazon | Amazonbot | Amzn-SearchBot (Alexa and other search experiences) |
| Mistral AI | MistralAI-Training | MistralAI-Index |
Common Crawl’s CCBot sits apart: it builds an open web archive that anyone can download and reuse, which is why many sites that block training crawlers block it too. The “Block AI training” preset includes it.
Two robots.txt rules break hand-written files without any error message, and the generator is built around both:
* group. If you write a group for GPTBot, the paths you disallowed for everyone no longer apply to GPTBot. So the generator only ever names a crawler to block it completely; everything else inherits the shared path rules.Google’s AdsBot also ignores the * group entirely. The “Block everything” preset names it, because a staging site that blocks * alone is still open to it.
noindex.https://www.example.com/robots.txt. It only applies to the protocol, host and port it is served from, so a subdomain needs its own.