How to appear in ChatGPT

AI search 21 September 2026·9 min read

There is no ranking in ChatGPT to climb, and no setting that puts you in the answer. What exists is plumbing: each engine runs two or three crawlers, and each one controls something different. Block the wrong one and you are not in the answer at all — which, on 21 September 2026, is the position a tenth of the sites I checked had put themselves in, mostly on purpose.

The three doors

Every AI engine that cites the web does three separate things, and gives each one its own crawler:

  • It builds a search index, so it has something to retrieve when a question arrives.
  • It fetches a page live, when a user pastes a link or asks about a specific site.
  • It collects training data for future models.

Those are different crawlers with different names, and the documentation is unusually clear about what each one costs you:

EngineSearch indexLive fetchTraining
ChatGPTOAI-SearchBotChatGPT-UserGPTBot
PerplexityPerplexityBotPerplexity-User—
ClaudeClaude-SearchBotClaude-UserClaudeBot
Google AI Overviews and AI ModeGooglebot—Google-Extended, for other systems

The search crawler is the one that decides whether you exist. OpenAI says OAI-SearchBot is “used to surface websites in search results in ChatGPT’s search features”, and that “sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.” Anthropic says blocking Claude-SearchBot “prevents our system from indexing your content for search optimization”. Perplexity says PerplexityBot is “designed to surface and link websites in search results on Perplexity” and, pointedly, that “it is not used to crawl content for AI foundation models.”

The training crawler is a licensing decision, not a visibility one. Blocking GPTBot or ClaudeBot keeps your pages out of future model weights. Anthropic describes a block as a signal that “the site’s future materials should be excluded from our AI model training datasets”. Neither statement says anything about whether you can be cited today, and the search crawlers are separate.

A grid of three AI engines against three jobs. ChatGPT uses OAI-SearchBot for its search index, ChatGPT-User for live fetches and GPTBot for training. Perplexity uses PerplexityBot and Perplexity-User, with no training crawler. Claude uses Claude-SearchBot, Claude-User and ClaudeBot. Two notes state that blocking the search crawler removes you from the answer, quoting OpenAI that opted-out sites will not be shown in ChatGPT search answers, and that blocking the training crawler is a licensing decision rather than a visibility one.
Three jobs, three crawlers, three different costs. The confusion that removes sites from AI answers is treating them as one decision.

Google works differently, and people get this backwards

AI Overviews and AI Mode are part of Search. They run on the Search index, which means Googlebot, and Google states plainly that “there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” Being indexed and eligible for snippets is the requirement.

Google-Extended is not the control people assume. Google describes it as limiting “AI training and grounding in some of Google’s other systems” — Gemini, in practice — and it does not govern whether you appear in AI Overviews or AI Mode.

The control that does govern that is newer: the Search generative AI setting in Search Console, rolled out to every site on 31 August 2026. Opt out and “links to your site and your site’s content won’t appear in Search generative AI features”, while, Google says, the control “isn’t used as a ranking or inclusion signal affecting other parts of Search.” It also “doesn’t affect AI training”, which stays with Google-Extended. Three separate switches, three separate consequences — and nosnippet, data-nosnippet and max-snippet still limit how much of a page can be shown.

What 230 brand sites actually allow

On 21 September 2026 I requested /robots.txt and /llms.txt from 230 sites — international retailers, SaaS, travel, banking, logistics, carmakers and twenty news publishers — and ran each file through the engine behind this site’s robots.txt tester, asking, for each of 28 crawlers, whether it may fetch the home page.

Sixty-two files could not be read at all: 32 answered 403, fourteen timed out, and the rest returned HTML, a 404, a 429 or a bot challenge. That is a finding in itself — a firewall that refuses a crawler makes the robots.txt policy irrelevant, because the fetch never gets that far. It left 168 readable files.

Of 168 readable robots.txt filesSitesShare
Name no AI crawler at all11769.6%
Block at least one training crawler3319.6%
Block at least one AI search crawler2313.7%
Block at least one user-triggered fetcher2112.5%
Block GPTBot2213.1%
Block OAI-SearchBot106.0%
Block PerplexityBot1810.7%
Block Googlebot00%

Three things stand out.

Most sites have no AI policy at all. Seven in ten readable files never name an AI crawler, so every one of them follows whatever the * group says. That is a decision by default, and for most businesses it happens to be the right one.

The split between training and retrieval is real, and deliberate. Of the 22 sites blocking GPTBot, twelve still allow OAI-SearchBot: no training, yes citation. Fourteen sites block a training crawler while allowing every AI search crawler — Zalando, Canva, Notion, Figma, Tripadvisor, Skyscanner, Babbel, Coursera, Netflix, SAP and Medium among them. That is the configuration most brands say they want when I ask, and it is straightforward to get right.

The ones who block everything are publishers. All ten sites blocking OAI-SearchBot are news organisations: the New York Times, the BBC, the Economist, El Mundo, ABC, Der Spiegel, FAZ, NRC, De Telegraaf and Dagens Nyheter. Across the twenty publishers in the sample, 100% block a training crawler and 90% block the AI search crawlers and the user-triggered fetchers. Among the 148 non-publishers, those figures are 9%, 1% and 1%. Publishers are negotiating licences; everyone else is trying to be found.

For scale, Ahrefs checked around 140 million websites in May 2025 and found GPTBot the most blocked AI bot at 5.89% of sites, with PerplexityBot at 5.61% and OAI-SearchBot at 5.54%. My sample blocks at roughly double those rates, which is what you would expect from large brands and newspapers rather than the web at large.

One more, for the llms.txt debate: 59 of the 230 sites (25.7%) serve a real llms.txt — a markdown file starting with a heading, rather than a catch-all HTML page. Adoption is no longer marginal. Whether anything reads it is a different question.

Blocking the user fetchers mostly does not do what people think

Eleven sites block ChatGPT-User and thirteen block Perplexity-User. Both vendors say those rules may not be honoured, because the request comes from a person, not a crawl: OpenAI writes that “because these actions are initiated by a user, robots.txt rules may not apply”, and Perplexity that “since a user requested the fetch, this fetcher generally ignores robots.txt rules.”

Anthropic is the exception that proves the point: it does honour a block on Claude-User, and says doing so “prevents our system from retrieving your content in response to a user query.” So the rule that works is the one that hurts — someone asked Claude to read your page, and Claude cannot.

Access makes you eligible, not cited

Nobody outside these companies knows how sources are weighted, and anyone selling you the weighting is guessing. What exists is correlation. Ahrefs looked at 75,000 brands in May 2025 and found brand web mentions correlated with AI Overview visibility at 0.664, against 0.218 for backlinks, 0.392 for branded search volume and 0.326 for Domain Rating. Twenty-six per cent of those brands never appeared at all. That study measures Google’s AI Overviews rather than ChatGPT, which is worth holding on to, because it gets quoted as though it covered everything.

Studies aimed specifically at ChatGPT contradict each other — some find it leans on brand sites, others on media and forums — and the reason is mundane: change the question set and you change the answer. The one finding that survives every sample is the unglamorous one, that being talked about elsewhere matters more than what you publish about yourself.

One thing that rarely gets measured: these engines answer the same question differently in different languages, and the sources cited for a Spanish query need not be the ones cited in English. Visibility in English tells you nothing about visibility in Spanish, so if you sell in both, measure both — with questions written in each language, and by country where it matters.

What to actually do

  1. Test what the engines can fetch

    Put your domain through the robots.txt tester. It answers per crawler, so you can see OAI-SearchBot, PerplexityBot and Claude-SearchBot separately from GPTBot and ClaudeBot.

  2. Check the firewall, not just the file

    Nearly a third of the sites I tried refused the request before robots.txt mattered. Ask whoever runs your WAF or bot protection whether the AI crawlers are allowed, and check their published IP ranges.

  3. Decide training separately from retrieval

    If legal wants out of training, block GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended — and leave the search crawlers alone. The robots.txt generator writes that file and checks it afterwards.

  4. Leave Googlebot alone

    AI Overviews and AI Mode come with the Search index. Use the Search Console control only if you genuinely want out of those features, and know it does not touch training.

  5. Measure what happens in Search

    Search Console’s generative AI reports give impressions by page and country, with no clicks and no queries. What that report can and cannot tell you is worth reading before you build a dashboard on it.

  6. Then, and only then, work on the answer

    Access gets you eligible. Being the page an answer needs is a content problem: answer engine optimisation for multi-market brands covers how engines pick sources, and why the smallest market is often the easiest to win.

What I would not spend time on

llms.txt as a visibility tactic. Publish one if you want a tidy fact sheet; the evidence that anything fetches it is still missing.

Schema markup as a magic ingredient. Google says there are no special requirements for AI features. Structured data earns its keep in ordinary Search; it is not a lever for AI Overviews.

Prompt-stuffing your own brand. Asking an engine about yourself until it says something nice is not measurement. If you want a baseline, ask real buyer questions across engines and count who gets cited — which is what the AI visibility checker does.

Sources

Questions

There is no ranking to climb. ChatGPT searches, retrieves a handful of pages and cites some of them. What you control is whether its crawlers may fetch you, whether the page answers the question, and whether the rest of the web makes you an obvious source.
No. GPTBot collects training data. The crawler that decides whether you appear is OAI-SearchBot, and OpenAI says sites opted out of it will not be shown in ChatGPT search answers. Most sites I checked that block training still allow it.
No. It limits AI training and grounding in other Google systems. AI Overviews and AI Mode run on the Search index, so Googlebot is the crawler, and the opt-out is the Search generative AI control in Search Console, worldwide since 31 August 2026.
Not reliably. OpenAI says robots.txt may not apply to ChatGPT-User because a person asked, and Perplexity says its user fetcher generally ignores it. Anthropic does honour a block on Claude-User — which means Claude cannot read the page someone asked it about.
No evidence that it does. A quarter of the sites I checked publish one, but server logs say almost nothing fetches it, and Google says its Search generative AI features do not use it. Fix crawler access first.