robots.txt generator

Choose which crawlers to keep out and which paths no one should crawl. The generator writes the file, then checks it crawler by crawler with Google’s matching rules — so blocking AI training does not quietly take you out of AI search.

Free, no email28 crawlersOnly usage recorded

Start from

Presets set the crawler list below. Change any box afterwards.

Crawlers to block

Ticked means blocked from the whole site. Hover a name for what blocking it costs.

Paths every other crawler should skip

One per line, starting with /. * matches anything; $ ends the URL.

robots.txt

Download

Upload it to the root of each host as /robots.txt. To test the live file afterwards, use the robots.txt tester.

Before you publish

What this file costs, in visibility.

    What the file does

    Checked, not assumed: every crawler run against the file above with the tester’s engine.

    Blocking AI training without leaving AI search

    Every large AI company now runs more than one crawler, and they do different jobs. The training crawlers collect content for future models. The search crawlers put your pages into answers that link back. Blocking the second kind to stop the first costs you visibility and protects nothing.

    CompanyTraining — safe to blockSearch and answers — blocking costs visibility
    GoogleGoogle-Extended (a token, not a crawler)Googlebot, which also covers AI Overviews and AI Mode
    OpenAIGPTBotOAI-SearchBot (ChatGPT search)
    AnthropicClaudeBotClaude-SearchBot, Claude-User
    AppleApplebot-Extended (a token, not a crawler)Applebot (Spotlight, Siri, Safari)
    Perplexity—PerplexityBot, which Perplexity says is not used for training
    MetaMeta-ExternalAgentMeta-WebIndexer (Meta AI search)
    AmazonAmazonbotAmzn-SearchBot (Alexa and other search experiences)
    Mistral AIMistralAI-TrainingMistralAI-Index

    Common Crawl’s CCBot sits apart: it builds an open web archive that anyone can download and reuse, which is why many sites that block training crawlers block it too. The “Block AI training” preset includes it.

    Why the generator checks its own output

    Two robots.txt rules break hand-written files without any error message, and the generator is built around both:

    • A named crawler ignores the * group. If you write a group for GPTBot, the paths you disallowed for everyone no longer apply to GPTBot. So the generator only ever names a crawler to block it completely; everything else inherits the shared path rules.
    • Some crawlers borrow another crawler’s group. Applebot follows Googlebot’s rules when no group names Applebot, and Googlebot-Image and Googlebot-News do the same. Block Googlebot and all three go with it — so if they should stay allowed, the generator gives them their own group with the shared rules copied in.

    Google’s AdsBot also ignores the * group entirely. The “Block everything” preset names it, because a staging site that blocks * alone is still open to it.

    What robots.txt cannot do

    • Keep a page out of the index. A blocked URL can still be indexed if other pages link to it; Google just cannot read it. To keep a page out, allow crawling and use noindex.
    • Stop crawlers that ignore it. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it, because a person asked for the page. A firewall rule is the reliable way to block those.
    • Remove you from AI Overviews alone. They are part of Google Search. Since 31 August 2026, Search Console’s Search generative AI control excludes a site from them without blocking Googlebot.
    • Protect anything private. The file is public and advisory. Anything that must stay private needs authentication.

    Sources

    Questions

    Block the training crawlers and tokens by name — GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, Amazonbot, CCBot — and leave the search crawlers alone: Googlebot, Bingbot, Applebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot. The “Block AI training” preset does exactly that.
    Not without leaving Google Search: AI Overviews and AI Mode are part of Search and use Googlebot. Since 31 August 2026, Search Console’s Search generative AI control excludes a site from them without blocking Googlebot. Google-Extended does not affect Search at all.
    At the root of each host, as https://www.example.com/robots.txt. It only applies to the protocol, host and port it is served from, so a subdomain needs its own.
    No. The file is built in your browser and the paths you type stay there. The page reports which preset you picked, which crawlers you blocked, and how many rules and sitemaps the file has — with no identifier attached. Do Not Track or Global Privacy Control stops even that.