11 tokens, 5 vendors, three different jobs — and blocking the wrong one quietly removes you from an assistant’s answers.
There are 11 AI crawler user agents and robots.txt control tokens worth knowing about, from 5 vendors, and they do three different jobs. 5 of them build a search index — block one and the assistant behind it can no longer find your pages. 3 fetch a single page while a user is mid-question. 2 collect training data, and blocking those does not remove you from anybody’s search results. One entry in the list is not a crawler at all. Every row below was checked against the vendor’s own bot documentation on 15 August 2026, and the doc link in each row is the page it was checked against.
| User agent token | Job | What it means for you | Source |
|---|---|---|---|
OAI-SearchBotOpenAI | search index | Indexes pages for ChatGPT’s search. Block it and ChatGPT’s search cannot reach you. | OpenAI doc |
ChatGPT-UserOpenAI | on-demand | Fetches one page while a user is asking, not a bulk crawl. | OpenAI doc |
GPTBotOpenAI | training | Collects training data. Blocking it does not remove you from ChatGPT’s search. | OpenAI doc |
PerplexityBotPerplexity | search index | Indexes pages for Perplexity’s results. Block it and Perplexity cannot reach you. | Perplexity doc |
Perplexity-UserPerplexity | on-demand | Fetches one page while a user is asking. | Perplexity doc |
Claude-SearchBotAnthropic | search index | Crawls to improve the search results Claude works from. | Anthropic doc |
Claude-UserAnthropic | on-demand | Fetches one page while a user is asking. | Anthropic doc |
ClaudeBotAnthropic | training | Collects training data. | Anthropic doc |
Googlebot | search index | Google’s index — AI Overviews are built on top of it. | Google doc |
Google-Extended | robots control | Not a crawler and not a user agent: a robots.txt token that governs Gemini training and grounding. It does not affect Google Search indexing or ranking. | Google doc |
bingbotMicrosoft | search index | Bing’s index. | Microsoft doc |
Vendors: Anthropic, Google, Microsoft, OpenAI, Perplexity. Rendered from the probe table the audit tool itself uses, so this page cannot drift from the tool.
The distinction that most robots.txt advice misses: an assistant’s search crawler and its training crawler are different tokens with different consequences.
OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, bingbot). These decide
whether the assistant can retrieve your page while answering. Disallow one and
you are not in that assistant’s answers, no matter how good the page is.ChatGPT-User, Perplexity-User, Claude-User). These arrive
because a user asked something that needs your page right now. Blocking them
breaks the live look-up but leaves the index alone.GPTBot, ClaudeBot). These feed model
training. Many sites block them deliberately, and that is a legitimate choice
with no cost to search visibility.GPTBot is OpenAI’s training crawler. A
User-agent: GPTBot / Disallow: / block keeps your pages out of
training data and does nothing to your visibility inside ChatGPT’s search
— that job belongs to OAI-SearchBot. So a rule meant to opt
out of training that disallows both tokens does two things, and only one of them
was intended: it opts you out of training, and it takes you out of the answers.
Check which tokens your own robots.txt names before you assume it
does what you meant.
The same shape repeats across vendors: ClaudeBot is training,
Claude-SearchBot is search. Google-Extended is the odd
one — Google’s own documentation describes it as a robots.txt control
token rather than a crawler, and it publishes no user-agent string at all.
Copy this, keep the comments, and change it deliberately rather than by pattern-matching someone’s blog post:
# The crawlers that decide whether an assistant can find you at all
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: bingbot
Allow: /
# Training crawlers — block these if that is your choice
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Two limits worth stating. robots.txt is voluntary
— it is a request, not an access control; anything you need actually
protected needs authentication. And a Disallow on
Google-Extended has no effect on Google Search indexing or ranking,
which is a separate system.
Some firewalls match on the whole user-agent string rather than the token, so sending only the token can produce a pass that a real crawler would not get. These are the strings the vendors publish:
| Token | User-agent string |
|---|---|
OAI-SearchBot | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot |
ChatGPT-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot |
GPTBot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot |
PerplexityBot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
Perplexity-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) |
Claude-SearchBot | Claude-SearchBot |
Claude-User | Claude-User |
ClaudeBot | ClaudeBot |
Googlebot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36 |
Google-Extended | — none published; this token only exists in robots.txt |
bingbot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36 |
To see what your own server does with one of them, send the request yourself:
curl -sI -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot' https://your-site.com/
A 200 means that crawler can read the page. A
403 — the usual symptom of a WAF or bot-protection rule, not of
robots.txt — means it cannot, and no amount of content work will fix that.
Claude-SearchBot, Claude-User and ClaudeBot
by name without publishing a full browser-style user-agent string, so that is
what the table shows — the token, not a reconstruction of it.Google-Extended as a crawler is the single most common error in
published AI-crawler lists.bingbot is in the scored set — but no vendor document we
checked states it, so treat it as a working assumption rather than a fact, and
weigh bingbot accordingly.If you want the same check run against your own site, with the reachability of each of these crawlers recorded as a scored audit line, that is one dimension of the $390 audit. Two of the assistants above are wired up today; three are not, and the home page keeps that list current.
The search-index ones: OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, bingbot. Each of those decides whether its assistant can retrieve your page while answering a question. Blocking a training crawler such as GPTBot or ClaudeBot has no effect on that.
No. GPTBot is OpenAI's training crawler. The crawler that indexes pages for ChatGPT's search feature is OAI-SearchBot, and it is a separate token you have to disallow separately.
No. It is a robots.txt control token that governs whether your content can be used for Gemini training and grounding, and Google publishes no user-agent string for it. Disallowing it does not affect Google Search indexing or ranking.
Send a request with that crawler's full user-agent string and read the status code: 200 means it gets the page, 403 usually means a firewall or bot-protection rule is refusing it before robots.txt is even consulted.
Every row was checked against the vendor's own bot documentation on 15 August 2026, and the doc link in each row is the page it was checked against. The table is generated from the probe table our audit tool uses, so it changes when the tool does.
geo_radar/probes.py, which is where the rows above are rendered from — each entry carries the doc URL it was checked against and the date it was checked.