There is no ranking in ChatGPT to climb, and no setting that puts you in the answer. What exists is plumbing: each engine runs two or three crawlers, and each one controls something different. Block the wrong one and you are not in the answer at all — which, on 21 September 2026, is the position a tenth of the sites I checked had put themselves in, mostly on purpose.
The three doors
Every AI engine that cites the web does three separate things, and gives each one its own crawler:
- It builds a search index, so it has something to retrieve when a question arrives.
- It fetches a page live, when a user pastes a link or asks about a specific site.
- It collects training data for future models.
Those are different crawlers with different names, and the documentation is unusually clear about what each one costs you:
| Engine | Search index | Live fetch | Training |
|---|---|---|---|
| ChatGPT | OAI-SearchBot | ChatGPT-User | GPTBot |
| Perplexity | PerplexityBot | Perplexity-User | — |
| Claude | Claude-SearchBot | Claude-User | ClaudeBot |
| Google AI Overviews and AI Mode | Googlebot | — | Google-Extended, for other systems |
The search crawler is the one that decides whether you exist. OpenAI says OAI-SearchBot is “used to surface websites in search results in ChatGPT’s search features”, and that “sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.” Anthropic says blocking Claude-SearchBot “prevents our system from indexing your content for search optimization”. Perplexity says PerplexityBot is “designed to surface and link websites in search results on Perplexity” and, pointedly, that “it is not used to crawl content for AI foundation models.”
The training crawler is a licensing decision, not a visibility one. Blocking GPTBot or ClaudeBot keeps your pages out of future model weights. Anthropic describes a block as a signal that “the site’s future materials should be excluded from our AI model training datasets”. Neither statement says anything about whether you can be cited today, and the search crawlers are separate.
Google works differently, and people get this backwards
AI Overviews and AI Mode are part of Search. They run on the Search index, which means Googlebot, and Google states plainly that “there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” Being indexed and eligible for snippets is the requirement.
Google-Extended is not the control people assume. Google describes it as limiting “AI training and grounding in some of Google’s other systems” — Gemini, in practice — and it does not govern whether you appear in AI Overviews or AI Mode.
The control that does govern that is newer: the Search generative AI setting in Search Console, rolled out to every site on 31 August 2026. Opt out and “links to your site and your site’s content won’t appear in Search generative AI features”, while, Google says, the control “isn’t used as a ranking or inclusion signal affecting other parts of Search.” It also “doesn’t affect AI training”, which stays with Google-Extended. Three separate switches, three separate consequences — and nosnippet, data-nosnippet and max-snippet still limit how much of a page can be shown.
What 230 brand sites actually allow
On 21 September 2026 I requested /robots.txt and /llms.txt from 230 sites — international retailers, SaaS, travel, banking, logistics, carmakers and twenty news publishers — and ran each file through the engine behind this site’s robots.txt tester, asking, for each of 28 crawlers, whether it may fetch the home page.
Sixty-two files could not be read at all: 32 answered 403, fourteen timed out, and the rest returned HTML, a 404, a 429 or a bot challenge. That is a finding in itself — a firewall that refuses a crawler makes the robots.txt policy irrelevant, because the fetch never gets that far. It left 168 readable files.
| Of 168 readable robots.txt files | Sites | Share |
|---|---|---|
| Name no AI crawler at all | 117 | 69.6% |
| Block at least one training crawler | 33 | 19.6% |
| Block at least one AI search crawler | 23 | 13.7% |
| Block at least one user-triggered fetcher | 21 | 12.5% |
Block GPTBot | 22 | 13.1% |
Block OAI-SearchBot | 10 | 6.0% |
Block PerplexityBot | 18 | 10.7% |
Block Googlebot | 0 | 0% |
Three things stand out.
Most sites have no AI policy at all. Seven in ten readable files never name an AI crawler, so every one of them follows whatever the * group says. That is a decision by default, and for most businesses it happens to be the right one.
The split between training and retrieval is real, and deliberate. Of the 22 sites blocking GPTBot, twelve still allow OAI-SearchBot: no training, yes citation. Fourteen sites block a training crawler while allowing every AI search crawler — Zalando, Canva, Notion, Figma, Tripadvisor, Skyscanner, Babbel, Coursera, Netflix, SAP and Medium among them. That is the configuration most brands say they want when I ask, and it is straightforward to get right.
The ones who block everything are publishers. All ten sites blocking OAI-SearchBot are news organisations: the New York Times, the BBC, the Economist, El Mundo, ABC, Der Spiegel, FAZ, NRC, De Telegraaf and Dagens Nyheter. Across the twenty publishers in the sample, 100% block a training crawler and 90% block the AI search crawlers and the user-triggered fetchers. Among the 148 non-publishers, those figures are 9%, 1% and 1%. Publishers are negotiating licences; everyone else is trying to be found.
For scale, Ahrefs checked around 140 million websites in May 2025 and found GPTBot the most blocked AI bot at 5.89% of sites, with PerplexityBot at 5.61% and OAI-SearchBot at 5.54%. My sample blocks at roughly double those rates, which is what you would expect from large brands and newspapers rather than the web at large.
One more, for the llms.txt debate: 59 of the 230 sites (25.7%) serve a real llms.txt — a markdown file starting with a heading, rather than a catch-all HTML page. Adoption is no longer marginal. Whether anything reads it is a different question.
Blocking the user fetchers mostly does not do what people think
Eleven sites block ChatGPT-User and thirteen block Perplexity-User. Both vendors say those rules may not be honoured, because the request comes from a person, not a crawl: OpenAI writes that “because these actions are initiated by a user, robots.txt rules may not apply”, and Perplexity that “since a user requested the fetch, this fetcher generally ignores robots.txt rules.”
Anthropic is the exception that proves the point: it does honour a block on Claude-User, and says doing so “prevents our system from retrieving your content in response to a user query.” So the rule that works is the one that hurts — someone asked Claude to read your page, and Claude cannot.
Access makes you eligible, not cited
Nobody outside these companies knows how sources are weighted, and anyone selling you the weighting is guessing. What exists is correlation. Ahrefs looked at 75,000 brands in May 2025 and found brand web mentions correlated with AI Overview visibility at 0.664, against 0.218 for backlinks, 0.392 for branded search volume and 0.326 for Domain Rating. Twenty-six per cent of those brands never appeared at all. That study measures Google’s AI Overviews rather than ChatGPT, which is worth holding on to, because it gets quoted as though it covered everything.
Studies aimed specifically at ChatGPT contradict each other — some find it leans on brand sites, others on media and forums — and the reason is mundane: change the question set and you change the answer. The one finding that survives every sample is the unglamorous one, that being talked about elsewhere matters more than what you publish about yourself.
One thing that rarely gets measured: these engines answer the same question differently in different languages, and the sources cited for a Spanish query need not be the ones cited in English. Visibility in English tells you nothing about visibility in Spanish, so if you sell in both, measure both — with questions written in each language, and by country where it matters.
What to actually do
Test what the engines can fetch
Put your domain through the robots.txt tester. It answers per crawler, so you can see
OAI-SearchBot,PerplexityBotandClaude-SearchBotseparately fromGPTBotandClaudeBot.Check the firewall, not just the file
Nearly a third of the sites I tried refused the request before robots.txt mattered. Ask whoever runs your WAF or bot protection whether the AI crawlers are allowed, and check their published IP ranges.
Decide training separately from retrieval
If legal wants out of training, block
GPTBot,ClaudeBot,CCBot,Google-ExtendedandApplebot-Extended— and leave the search crawlers alone. The robots.txt generator writes that file and checks it afterwards.Leave Googlebot alone
AI Overviews and AI Mode come with the Search index. Use the Search Console control only if you genuinely want out of those features, and know it does not touch training.
Measure what happens in Search
Search Console’s generative AI reports give impressions by page and country, with no clicks and no queries. What that report can and cannot tell you is worth reading before you build a dashboard on it.
Then, and only then, work on the answer
Access gets you eligible. Being the page an answer needs is a content problem: answer engine optimisation for multi-market brands covers how engines pick sources, and why the smallest market is often the easiest to win.
What I would not spend time on
llms.txt as a visibility tactic. Publish one if you want a tidy fact sheet; the evidence that anything fetches it is still missing.
Schema markup as a magic ingredient. Google says there are no special requirements for AI features. Structured data earns its keep in ordinary Search; it is not a lever for AI Overviews.
Prompt-stuffing your own brand. Asking an engine about yourself until it says something nice is not measurement. If you want a baseline, ask real buyer questions across engines and count who gets cited — which is what the AI visibility checker does.
Sources
- OpenAI, Bots and crawlers:
OAI-SearchBot,GPTBotandChatGPT-User. - Anthropic, Does Anthropic crawl data from the web?:
ClaudeBot,Claude-SearchBotandClaude-User. - Perplexity, PerplexityBot and Perplexity-User.
- Google Search Central, AI features and your website, last updated 10 December 2025, and Search Console Help, Search generative AI control, rolled out worldwide on 31 August 2026.
- Patrick Stox and Xibeijia Guan, The AI bots that ~140 million websites block the most, Ahrefs, 21 May 2025.
- Louise Linehan, An analysis of AI Overview brand visibility factors, Ahrefs, 26 May 2025, on 75,000 brands.
/robots.txtand/llms.txtof 230 international brand, SaaS and publisher sites, requested once each on 21 September 2026 with this site’s robots.txt fetcher: 168 readable files, evaluated for 28 crawlers with the tester’s own engine.
Questions
GPTBot collects training data. The crawler that decides whether you appear is OAI-SearchBot, and OpenAI says sites opted out of it will not be shown in ChatGPT search answers. Most sites I checked that block training still allow it.ChatGPT-User because a person asked, and Perplexity says its user fetcher generally ignores it. Anthropic does honour a block on Claude-User — which means Claude cannot read the page someone asked it about.