robots.txt tester

Is a URL blocked, for which crawler, and by which rule? Test Googlebot, Bingbot and the AI crawlers — GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and more — against a live or pasted file. It applies Google’s matching rules, including the ones that catch people out.

Free, no email28 crawlersAddress recorded, not files

Page to test

Any URL on the site. The tester fetches that site’s /robots.txt and checks the URL against it.

Crawlers

Hover a name for what it does. The common set covers Google, Bing, Apple, OpenAI, Anthropic and Perplexity.

How the tester decides

It follows Google’s documentation of robots.txt and the open-source parser that documentation describes. Four rules do almost all of the work:

  1. A crawler obeys one group

    The group whose user agent matches it most specifically. If a group names the crawler, the * group no longer applies to it at all — not even the rules you meant for everyone. Several groups for the same crawler are combined.

  2. Blank lines do not end a group

    Consecutive user-agent lines share the rules that follow, even with a blank line or a comment between them. Only a rule ends a group.

  3. The longest matching rule wins

    Measured by the length of the rule’s path. On a tie, allow wins. * matches any run of characters and $ marks the end of the URL. Paths are case-sensitive and include the query string.

  4. Only four fields count for Google

    user-agent, allow, disallow and sitemap. crawl-delay and noindex are ignored, and anything after the first 500 KiB of the file is too.

Some crawlers add their own twist, and the tester models each one from its owner’s documentation. Googlebot-Image and Googlebot-News follow the Googlebot group when they have none of their own. Applebot follows Googlebot’s rules when no group names Applebot. And Google’s AdsBot ignores the * group, so it is only blocked when it is named.

Search, AI search and AI training are different bots

Most confusion about AI crawlers comes from treating each company as one bot. They are not, and blocking the wrong one costs you visibility without protecting anything:

If you want to…The crawler that decidesNot this one
Stay out of Google’s AI Overviews and AI ModeNo crawler: the Search generative AI control in Search Console, available to every site since 31 August 2026. Blocking Googlebot would remove you from Search altogetherGoogle-Extended, which covers Gemini training and grounding and does not affect Search
Keep content out of OpenAI’s trainingGPTBotOAI-SearchBot, which puts you in ChatGPT search
Appear in ChatGPT search answersOAI-SearchBot must be allowedGPTBot
Keep content out of Anthropic’s trainingClaudeBotClaude-SearchBot and Claude-User, which affect visibility in Claude
Keep content out of Apple’s AI trainingApplebot-ExtendedApplebot, which powers Spotlight, Siri and Safari search

One more caution. Fetchers that act on a user’s request may not follow robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores it, and Meta and Amazon say something similar about their user-triggered fetchers. Anthropic says Claude-User respects it. For the ones that may not, the tester still shows what the file says, and flags the gap.

What the status code means

For a live site the tester reports what the server answered, and explains it the way Google does:

  • 200. The file is read and applied.
  • Redirects. Google follows at least five hops. Beyond that it treats robots.txt as not found.
  • 4xx, except 429. Treated as if there were no robots.txt: nothing is blocked. A 403 on robots.txt does not keep Google out.
  • 429, 5xx or a network error. Google stops crawling the site for 12 hours while it retries, then uses the last copy it could read for up to 30 days.
  • An HTML page served as robots.txt. Google parses it anyway, finds no valid rules and blocks nothing — usually not what the site intended.

What this tester sends

A pasted file never leaves your browser. For a live site, the address you enter is sent to a function on this site, which requests /robots.txt from that host — and nothing else — following redirects with the same checks as the hreflang checker. It returns the status and the rule lines, never the rest of the response, and stores nothing. A site can serve this tester a different file than it serves Googlebot; if you suspect that, compare with the robots.txt report in Search Console.

To write a new file rather than test one, the robots.txt generator builds one with the same rules, and checks its own output with this tester’s engine.

Sources

Questions

No. In November 2023 Google added a robots.txt report to Search Console and retired the old Tester. The report shows the files Google found and their errors, and URL Inspection shows whether a page is blocked for Googlebot — but neither tests a rule against a URL for other crawlers.
No. GPTBot is OpenAI’s training crawler. ChatGPT search relies on OAI-SearchBot, and OpenAI says sites that block it will not be shown in ChatGPT search answers. The two are controlled separately.
No. Google says it does not affect inclusion in Search and is not a ranking signal. It controls Gemini training and grounding. AI Overviews and AI Mode are part of Search, which Googlebot crawls.
A pasted file is tested in your browser and its contents are never sent. For a live site, the address goes to a function on this site that fetches that site’s robots.txt and returns its rule lines. The host is recorded, with the status and how many crawlers and paths you tested. Do Not Track or Global Privacy Control stops that record.