Is a URL blocked, for which crawler, and by which rule? Test Googlebot, Bingbot and the AI crawlers — GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and more — against a live or pasted file. It applies Google’s matching rules, including the ones that catch people out.
For the URL you entered. “Decided by” names the rule and the group it came from.
✓ allowed, ✗ blocked. Hover a cell for the deciding rule.
Problems first. The rules are Google’s unless a line says otherwise.
It follows Google’s documentation of robots.txt and the open-source parser that documentation describes. Four rules do almost all of the work:
The group whose user agent matches it most specifically. If a group names the crawler, the * group no longer applies to it at all — not even the rules you meant for everyone. Several groups for the same crawler are combined.
Consecutive user-agent lines share the rules that follow, even with a blank line or a comment between them. Only a rule ends a group.
Measured by the length of the rule’s path. On a tie, allow wins. * matches any run of characters and $ marks the end of the URL. Paths are case-sensitive and include the query string.
user-agent, allow, disallow and sitemap. crawl-delay and noindex are ignored, and anything after the first 500 KiB of the file is too.
Some crawlers add their own twist, and the tester models each one from its owner’s documentation. Googlebot-Image and Googlebot-News follow the Googlebot group when they have none of their own. Applebot follows Googlebot’s rules when no group names Applebot. And Google’s AdsBot ignores the * group, so it is only blocked when it is named.
Most confusion about AI crawlers comes from treating each company as one bot. They are not, and blocking the wrong one costs you visibility without protecting anything:
| If you want to… | The crawler that decides | Not this one |
|---|---|---|
| Stay out of Google’s AI Overviews and AI Mode | No crawler: the Search generative AI control in Search Console, available to every site since 31 August 2026. Blocking Googlebot would remove you from Search altogether | Google-Extended, which covers Gemini training and grounding and does not affect Search |
| Keep content out of OpenAI’s training | GPTBot | OAI-SearchBot, which puts you in ChatGPT search |
| Appear in ChatGPT search answers | OAI-SearchBot must be allowed | GPTBot |
| Keep content out of Anthropic’s training | ClaudeBot | Claude-SearchBot and Claude-User, which affect visibility in Claude |
| Keep content out of Apple’s AI training | Applebot-Extended | Applebot, which powers Spotlight, Siri and Safari search |
One more caution. Fetchers that act on a user’s request may not follow robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores it, and Meta and Amazon say something similar about their user-triggered fetchers. Anthropic says Claude-User respects it. For the ones that may not, the tester still shows what the file says, and flags the gap.
For a live site the tester reports what the server answered, and explains it the way Google does:
A pasted file never leaves your browser. For a live site, the address you enter is sent to a function on this site, which requests /robots.txt from that host — and nothing else — following redirects with the same checks as the hreflang checker. It returns the status and the rule lines, never the rest of the response, and stores nothing. A site can serve this tester a different file than it serves Googlebot; if you suspect that, compare with the robots.txt report in Search Console.
To write a new file rather than test one, the robots.txt generator builds one with the same rules, and checks its own output with this tester’s engine.
OAI-SearchBot, and OpenAI says sites that block it will not be shown in ChatGPT search answers. The two are controlled separately.