AI crawlers: the complete list and what each one does

Roughly half the sites we audit are blocking at least one major AI crawler. Almost none of them meant to.

It is invisible in analytics, it never appears in Search Console, and the site keeps ranking normally on Google throughout. The only symptom is absence: the business simply never gets named in AI answers, and nobody can work out why.

This page lists every AI crawler worth knowing, what each one actually does, and how to check in about ten minutes whether yours can be read.

Last updated · · Part of our AI search knowledge base
01The short version

The short version

If you want AI assistants to find, read and recommend your site, this is the robots.txt block you want.

robots.txt
# AI search and retrieval: allow
User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Googlebot
Allow: /

User-agent: bingbot
Allow: /

User-agent: Applebot
Allow: /

# AI training: your choice, see below
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: Applebot-Extended
Allow: /

Two things before you paste that anywhere.

A missing rule is not a block. Crawlers default to allowed. You only need explicit Allow lines if something upstream is disallowing them, or to make your intent obvious to whoever maintains the file next. Explicit is usually better than implicit here, because the person who breaks this in eight months will be reading your robots.txt, not your notes.

robots.txt is only half the story. Most accidental blocks happen at the CDN or firewall layer, where robots.txt has no authority at all. More on that below, and it is the section most people need.

02Three jobs, not one

Three jobs, not one

The single most common misunderstanding: treating “AI crawlers” as one category. They are not. Most major vendors run three separate bots doing three different jobs, and they can be controlled independently.

Job 01Training
GPTBotClaudeBotmeta-externalagentCCBot
What it does

Collects content that may be used to train future models

If you block itLow cost

Your content is excluded from future training data. Existing model knowledge is unaffected

Job 02Search indexing
OAI-SearchBotClaude-SearchBotPerplexityBotGooglebotbingbot
What it does

Builds the index the assistant retrieves from when answering

If you block itHigh cost

This is the expensive one. You are removed from the retrieval pool entirely

Job 03User-initiated fetch
ChatGPT-UserClaude-UserPerplexity-User
What it does

Pulls a specific page because a user asked a question that needs it

If you block itMedium cost

Live lookups of your pages fail

Each job is controlled separately in robots.txt

Get this distinction right and everything else on this page follows.

A publisher who wants their archive kept out of training datasets can block the training bot while leaving the search and user bots fully open. They stay visible in AI answers and opt out of training. That is a coherent, defensible position, and almost nobody realizes it is available, because the vendors present the bots as a list rather than as a choice.

The failure mode is the opposite: someone reads an article about AI training, adds a blanket disallow, and quietly removes their business from AI search results as a side effect they never intended.

03The complete list

The complete list

25 crawlers
TokenJobRespects robots.txt
OpenAI
GPTBotTrainingYes
OAI-SearchBotSearch indexingYes
ChatGPT-UserUser-initiated fetchYes, though OpenAI notes rules may not apply to user-triggered fetches
Anthropic
ClaudeBotTrainingYes
Claude-SearchBotSearch indexingYes
Claude-UserUser-initiated fetchYes

Anthropic is unusual in that all three honour robots.txt, including the user-initiated one. Do not assume that consistency across vendors.

Google
GooglebotSearch, and all Google AI surfacesYes
Google-ExtendedTraining opt-out control for GeminiNot a crawler, a permission token

The important point, and the one that surprises people: Google-Extended does not remove you from AI Overviews. It governs training and grounding improvement only. There is no way to appear in Google Search while opting out of AI Overviews, because Googlebot serves both.

Perplexity
PerplexityBotSearch indexingYes
Perplexity-UserLive retrieval for a user requestYes, with the usual user-initiated caveat
Microsoft
bingbotBing search, and Copilot groundingYes

Microsoft runs no separate AI crawler. Copilot grounds in Bing’s index through Prometheus, so Bingbot access is Copilot access. Since ChatGPT also leans heavily on Bing, this single token matters more than its plainness suggests.

Apple
ApplebotSiri, Spotlight, Apple searchYes
Applebot-ExtendedTraining opt-out control for Apple IntelligenceNot a crawler, a permission token
Meta
meta-externalagentTraining and content fetching for Meta AIYes
meta-externalfetcherReal-time fetching for Meta AIYes
FacebookBotMeta AI trainingYes
xAI (Grok)

xAI has not published crawler documentation comparable to OpenAI’s or Anthropic’s, and the tokens circulating online are inconsistent. Grok’s retrieval leans primarily on X rather than a broad conventional web crawl.

We deliberately are not listing a token here, because publishing an unverified string would send your robots.txt rule nowhere while giving you false confidence that it worked. If you need to write a rule for xAI, check its current documentation directly rather than copying a list, including this one.

Everyone else
CCBotCommon CrawlOpen dataset widely used for AI trainingVariable
AmazonbotAmazonAlexa and Amazon AIYes
BytespiderByteDanceTrainingKnown compliance issues
DiffbotDiffbotKnowledge graph buildingVariable
cohere-aiCohereTrainingYes
YouBotYou.comAI search indexingYes
omgili / omgilibotWebz.ioData aggregation for AIVariable
iaskspideriAsk.aiAI search indexingYes
img2datasetHugging FaceImage dataset collectionNo

The bottom four rows are where good intentions run out. For bots with poor compliance records, robots.txt is a request rather than a control. If you genuinely need to stop them, that happens at the server or CDN with a 403, not in a text file.

Tokens as published by each operator. Check vendor documentation before relying on any list, including this one
04The four that matter

The four that matter most

05Training vs search

Training vs search: the decision you actually have

Start here
What is your goal?
Most businesses
Maximum AI visibility
Training bots
Allow
Search bots
Allow
Publishers whose content is the product
Stay visible in answers, opt out of training
Training bots
Block
GPTBot, ClaudeBot, Google-Extended, Applebot-Extended
Search bots
Allow all
Rare. Think hard
Complete withdrawal from AI
Training bots
Block
Search bots
Block
The training and search decisions are separable

Most businesses should allow everything. If your goal is to be found and recommended, there is no upside to withholding content from training: it is the raw material from which models learn your brand exists.

Publishers are the real exception, and the calculation is genuinely different. If your content is the product, and an assistant summarizing it means nobody visits, then training exclusion is a rational commercial position rather than a reflex.

Here is the part that gets missed. Those two decisions are separable.

The middle route is the one almost nobody knows exists. You keep every citation opportunity and remove your content from future training datasets. For a publisher weighing this, it is usually the right answer, and it is available today with a handful of lines.

One caveat worth being straight about: blocking training bots now does nothing about models already trained. That content is in. This is a forward-looking control, not an eraser.

06Beyond robots.txt

Why robots.txt is not the whole story

This is the section most people need, and the reason accidental blocks are so common.

Your robots.txt can be perfect and your site can still be unreadable, because the request never reaches robots.txt in the first place.

RequestAI crawlerSends the request
403 hereCDN bot protectionBot-fight mode or an AI blocking toggle returns 403
Layer 02WAF rulesRate limits and bot heuristics
Layer 03Server and pluginsOld rules nobody remembers adding
Layer 04robots.txtNever consulted
Layer 05Your pageNever read

The request is refused before it reaches robots.txt. Your file can be perfect and the site still unreadable.

Nothing in your analytics or Search Console reports this

Where blocks actually happen:

LayerTypical causeVisible in robots.txt?
CDN bot protectionCloudflare bot-fight mode, AI crawler blocking toggles, often on by defaultNo
WAF rulesAggressive rate limiting or bot heuristicsNo
Security pluginsWordPress plugins blocking unfamiliar agents out of the boxNo
Server configOld .htaccess or nginx rules nobody remembers addingNo
JavaScript renderingContent that only exists after executionNot a block, but the same effect

That last row deserves emphasis. Most AI crawlers do not execute JavaScript. A client-side rendered site is not blocked, exactly. It is simply empty when read.

Server-side rendering is not a nice-to-have for AI visibility, it is the difference between having content and appearing to have none.

Several CDN providers added AI crawler blocking as a default or near-default setting during 2025 and 2026. A significant share of the accidental blocks we find trace back to a toggle nobody consciously flipped.

07How to check

How to actually check

Ten minutes, in this order.

01

Read your robots.txt properly

Visit yourdomain.com/robots.txt and look for any Disallow: / under an AI token, plus any blanket User-agent: * rule that catches everything.

Check every subdomain separately. Rules do not inherit. shop.yourdomain.com and blog.yourdomain.com each need their own file.

02

Check your CDN and firewall

In Cloudflare, look at Security, then Bots, and specifically at any AI crawler control. Other providers have equivalents under bot management. This is where the surprises are.

03

Read your server logs

The only reliable evidence. Everything else is what you have configured; logs are what actually happened.

Search the last 30 days for these tokens: GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, bingbot, Googlebot.

You are looking for two things: whether they visited at all, and what status codes they received. A pattern of 403s is a block. Silence is worse, because it usually means they never got that far.

04

Test your rendering

Fetch a key page with JavaScript disabled, or use a plain HTTP request rather than a browser. If the page comes back empty, that is what most AI crawlers see.

If steps 1 and 2 look fine but the logs show nothing, the problem is upstream of both. That is the moment to bring in someone who reads logs for a living. Once crawlers can reach you, the next question is whether AI actually names you, which is what measuring AI visibility answers.

08Should you block?

Should you block anything?

A balanced view, because the reflexive answers in both directions are wrong.

Reasons to allow everything:

  • You want to be found, cited and recommended
  • Training exposure is how models learn your brand exists at all
  • The cost of accidental over-blocking is far higher than the cost of being trained on

Reasons a publisher might block training bots:

  • Your content is the product and summarization substitutes for visiting
  • You are pursuing or maintaining a licensing position
  • You hold rights that make broad reuse genuinely problematic

Reasons almost nobody should block search bots:

  • It removes you from AI answers entirely, with no compensating benefit

Reasons blocking may not work anyway:

  • Several crawlers have poor compliance records
  • Some data reaches models through third-party datasets you never see
  • Content already used in training cannot be withdrawn

If you do block, block precisely. A blanket rule is how a business intending to protect its archive accidentally removes itself from ChatGPT.

09Seven mistakes

Seven common mistakes

The mistakeWhat it costsThe fix
Blocking GPTBot to “opt out of AI”Nothing gained on visibility, and it does not affect search retrieval anywayUnderstand the three jobs before writing any rule
Assuming robots.txt is the whole pictureSilent CDN block, invisible in every dashboardCheck CDN and firewall settings directly
Forgetting subdomainsWhole sections unreadableOne robots.txt per subdomain
Client-side rendering onlyCrawlers see an empty pageServer-side render
Blocking by IP rangeUnreliable, most vendors do not publish ranges, can block robots.txt itselfMatch on user agent instead
Copying a robots.txt from a blog postStale tokens that match nothingVerify tokens against vendor documentation
Never checking the logsNo idea whether any of it workedQuarterly log review

Structured data is the next layer down once crawlers can read you: see our guide to structured data for AI.

10FAQ

Frequently asked questions

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects content that may train future OpenAI models. OAI-SearchBot builds the index ChatGPT retrieves from when answering questions. They are independent: blocking GPTBot does not remove you from ChatGPT’s search results, and blocking OAI-SearchBot does. More on how ChatGPT chooses sources in our ChatGPT SEO guide.

Does Google-Extended stop me appearing in AI Overviews?

No. Google-Extended controls whether your content is used to train and ground Gemini models. AI Overviews are served from the standard Google index via Googlebot, so there is no way to stay in Google Search while opting out of AI Overviews.

How do I know if I am blocking AI crawlers?

Check robots.txt on every subdomain, then your CDN and firewall settings, then your server logs for the last 30 days. Logs are the only reliable evidence, because they show what actually happened rather than what you configured.

Do AI crawlers respect robots.txt?

The major vendors state that they do: OpenAI, Anthropic, Google, Perplexity, Microsoft, Apple and Meta. Several smaller crawlers have mixed or poor compliance records, including Bytespider and some dataset collectors. For those, enforcement requires a server-level block rather than a robots.txt rule.

Should I block AI crawlers?

Most businesses should not. If you want to be recommended by AI assistants, blocking removes you from consideration. Publishers whose content is the product have a genuine case for blocking training bots specifically, while leaving search bots open.

Can AI crawlers read JavaScript?

Mostly no. A site that renders content client-side may be technically accessible and still functionally empty to an AI crawler. Server-side rendering resolves it, and it’s how every website we build is served.

How often should I check this?

Quarterly, and after any CDN change, security plugin installation or site migration. Those three events cause most of the accidental blocks we find.

Does blocking a training bot remove my content from models already trained?

No. It is forward-looking only. Content already used in training cannot be withdrawn by a robots.txt rule.
Free AI visibility audit

Not sure what is reaching your site?

Crawler access is the first thing we check on any audit, because it is binary. If the models cannot read you, nothing else you do matters.

We check your robots.txt across every subdomain, your CDN and firewall rules, and 30 days of server logs, then show you exactly which crawlers reached you and what they got back.