AI crawler user agents

11 tokens, 5 vendors, three different jobs — and blocking the wrong one quietly removes you from an assistant’s answers.

The AI crawler user agents, and what each one is for

There are 11 AI crawler user agents and robots.txt control tokens worth knowing about, from 5 vendors, and they do three different jobs. 5 of them build a search index — block one and the assistant behind it can no longer find your pages. 3 fetch a single page while a user is mid-question. 2 collect training data, and blocking those does not remove you from anybody’s search results. One entry in the list is not a crawler at all. Every row below was checked against the vendor’s own bot documentation on 15 August 2026, and the doc link in each row is the page it was checked against.

User agent tokenJobWhat it means for youSource
OAI-SearchBot
OpenAI
search indexIndexes pages for ChatGPT’s search. Block it and ChatGPT’s search cannot reach you.OpenAI doc
ChatGPT-User
OpenAI
on-demandFetches one page while a user is asking, not a bulk crawl.OpenAI doc
GPTBot
OpenAI
trainingCollects training data. Blocking it does not remove you from ChatGPT’s search.OpenAI doc
PerplexityBot
Perplexity
search indexIndexes pages for Perplexity’s results. Block it and Perplexity cannot reach you.Perplexity doc
Perplexity-User
Perplexity
on-demandFetches one page while a user is asking.Perplexity doc
Claude-SearchBot
Anthropic
search indexCrawls to improve the search results Claude works from.Anthropic doc
Claude-User
Anthropic
on-demandFetches one page while a user is asking.Anthropic doc
ClaudeBot
Anthropic
trainingCollects training data.Anthropic doc
Googlebot
Google
search indexGoogle’s index — AI Overviews are built on top of it.Google doc
Google-Extended
Google
robots controlNot a crawler and not a user agent: a robots.txt token that governs Gemini training and grounding. It does not affect Google Search indexing or ranking.Google doc
bingbot
Microsoft
search indexBing’s index.Microsoft doc

Vendors: Anthropic, Google, Microsoft, OpenAI, Perplexity. Rendered from the probe table the audit tool itself uses, so this page cannot drift from the tool.

Blocking the wrong one costs you the answer, not the training set

The distinction that most robots.txt advice misses: an assistant’s search crawler and its training crawler are different tokens with different consequences.

The gptbot robots.txt rule almost everyone gets backwards

GPTBot is OpenAI’s training crawler. A User-agent: GPTBot / Disallow: / block keeps your pages out of training data and does nothing to your visibility inside ChatGPT’s search — that job belongs to OAI-SearchBot. So a rule meant to opt out of training that disallows both tokens does two things, and only one of them was intended: it opts you out of training, and it takes you out of the answers. Check which tokens your own robots.txt names before you assume it does what you meant.

The same shape repeats across vendors: ClaudeBot is training, Claude-SearchBot is search. Google-Extended is the odd one — Google’s own documentation describes it as a robots.txt control token rather than a crawler, and it publishes no user-agent string at all.

A robots.txt that blocks AI crawlers for training and keeps you findable

Copy this, keep the comments, and change it deliberately rather than by pattern-matching someone’s blog post:

# The crawlers that decide whether an assistant can find you at all
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: bingbot
Allow: /

# Training crawlers — block these if that is your choice
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Two limits worth stating. robots.txt is voluntary — it is a request, not an access control; anything you need actually protected needs authentication. And a Disallow on Google-Extended has no effect on Google Search indexing or ranking, which is a separate system.

The full user-agent strings, and how to test your own site

Some firewalls match on the whole user-agent string rather than the token, so sending only the token can produce a pass that a real crawler would not get. These are the strings the vendors publish:

TokenUser-agent string
OAI-SearchBotMozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
ChatGPT-UserMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
GPTBotMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
PerplexityBotMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity-UserMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
Claude-SearchBotClaude-SearchBot
Claude-UserClaude-User
ClaudeBotClaudeBot
GooglebotMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36
Google-Extended— none published; this token only exists in robots.txt
bingbotMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36

To see what your own server does with one of them, send the request yourself:

curl -sI -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot' https://your-site.com/

A 200 means that crawler can read the page. A 403 — the usual symptom of a WAF or bot-protection rule, not of robots.txt — means it cannot, and no amount of content work will fix that.

Three things this table does not pretend

If you want the same check run against your own site, with the reachability of each of these crawlers recorded as a scored audit line, that is one dimension of the $390 audit. Two of the assistants above are wired up today; three are not, and the home page keeps that list current.

AI crawler questions we get asked

Which AI crawler user agents should I never block?

The search-index ones: OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, bingbot. Each of those decides whether its assistant can retrieve your page while answering a question. Blocking a training crawler such as GPTBot or ClaudeBot has no effect on that.

Does blocking GPTBot in robots.txt hide me from ChatGPT?

No. GPTBot is OpenAI's training crawler. The crawler that indexes pages for ChatGPT's search feature is OAI-SearchBot, and it is a separate token you have to disallow separately.

Is Google-Extended a crawler?

No. It is a robots.txt control token that governs whether your content can be used for Gemini training and grounding, and Google publishes no user-agent string for it. Disallowing it does not affect Google Search indexing or ranking.

How do I check whether my site already blocks one of them?

Send a request with that crawler's full user-agent string and read the status code: 200 means it gets the page, 403 usually means a firewall or bot-protection rule is refusing it before robots.txt is even consulted.

How current is this list?

Every row was checked against the vendor's own bot documentation on 15 August 2026, and the doc link in each row is the page it was checked against. The table is generated from the probe table our audit tool uses, so it changes when the tool does.

Where these numbers come from