AI crawlers: the complete list and what each one does
Roughly half the sites we audit are blocking at least one major AI crawler. Almost none of them meant to.
It is invisible in analytics, it never appears in Search Console, and the site keeps ranking normally on Google throughout. The only symptom is absence: the business simply never gets named in AI answers, and nobody can work out why.
This page lists every AI crawler worth knowing, what each one actually does, and how to check in about ten minutes whether yours can be read.
The short version
If you want AI assistants to find, read and recommend your site, this is the robots.txt block you want.
# AI search and retrieval: allow
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: Googlebot
Allow: /
User-agent: bingbot
Allow: /
User-agent: Applebot
Allow: /
# AI training: your choice, see below
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot-Extended
Allow: /Two things before you paste that anywhere.
A missing rule is not a block. Crawlers default to allowed. You only need explicit Allow lines if something upstream is disallowing them, or to make your intent obvious to whoever maintains the file next. Explicit is usually better than implicit here, because the person who breaks this in eight months will be reading your robots.txt, not your notes.
robots.txt is only half the story. Most accidental blocks happen at the CDN or firewall layer, where robots.txt has no authority at all. More on that below, and it is the section most people need.
Three jobs, not one
The single most common misunderstanding: treating “AI crawlers” as one category. They are not. Most major vendors run three separate bots doing three different jobs, and they can be controlled independently.
Collects content that may be used to train future models
Your content is excluded from future training data. Existing model knowledge is unaffected
Builds the index the assistant retrieves from when answering
This is the expensive one. You are removed from the retrieval pool entirely
Pulls a specific page because a user asked a question that needs it
Live lookups of your pages fail
Get this distinction right and everything else on this page follows.
A publisher who wants their archive kept out of training datasets can block the training bot while leaving the search and user bots fully open. They stay visible in AI answers and opt out of training. That is a coherent, defensible position, and almost nobody realizes it is available, because the vendors present the bots as a list rather than as a choice.
The failure mode is the opposite: someone reads an article about AI training, adds a blanket disallow, and quietly removes their business from AI search results as a side effect they never intended.
The complete list
| Token | Job | Respects robots.txt |
|---|---|---|
| OpenAI | ||
| GPTBot | Training | Yes |
| OAI-SearchBot | Search indexing | Yes |
| ChatGPT-User | User-initiated fetch | Yes, though OpenAI notes rules may not apply to user-triggered fetches |
| Anthropic | ||
| ClaudeBot | Training | Yes |
| Claude-SearchBot | Search indexing | Yes |
| Claude-User | User-initiated fetch | Yes |
Anthropic is unusual in that all three honour robots.txt, including the user-initiated one. Do not assume that consistency across vendors. | ||
| Googlebot | Search, and all Google AI surfaces | Yes |
| Google-Extended | Training opt-out control for Gemini | Not a crawler, a permission token |
The important point, and the one that surprises people: Google-Extended does not remove you from AI Overviews. It governs training and grounding improvement only. There is no way to appear in Google Search while opting out of AI Overviews, because Googlebot serves both. | ||
| Perplexity | ||
| PerplexityBot | Search indexing | Yes |
| Perplexity-User | Live retrieval for a user request | Yes, with the usual user-initiated caveat |
| Microsoft | ||
| bingbot | Bing search, and Copilot grounding | Yes |
Microsoft runs no separate AI crawler. Copilot grounds in Bing’s index through Prometheus, so Bingbot access is Copilot access. Since ChatGPT also leans heavily on Bing, this single token matters more than its plainness suggests. | ||
| Apple | ||
| Applebot | Siri, Spotlight, Apple search | Yes |
| Applebot-Extended | Training opt-out control for Apple Intelligence | Not a crawler, a permission token |
| Meta | ||
| meta-externalagent | Training and content fetching for Meta AI | Yes |
| meta-externalfetcher | Real-time fetching for Meta AI | Yes |
| FacebookBot | Meta AI training | Yes |
| xAI (Grok) | ||
xAI has not published crawler documentation comparable to OpenAI’s or Anthropic’s, and the tokens circulating online are inconsistent. Grok’s retrieval leans primarily on X rather than a broad conventional web crawl. | ||
We deliberately are not listing a token here, because publishing an unverified string would send your robots.txt rule nowhere while giving you false confidence that it worked. If you need to write a rule for xAI, check its current documentation directly rather than copying a list, including this one. | ||
| Everyone else | ||
| CCBotCommon Crawl | Open dataset widely used for AI training | Variable |
| AmazonbotAmazon | Alexa and Amazon AI | Yes |
| BytespiderByteDance | Training | Known compliance issues |
| DiffbotDiffbot | Knowledge graph building | Variable |
| cohere-aiCohere | Training | Yes |
| YouBotYou.com | AI search indexing | Yes |
| omgili / omgilibotWebz.io | Data aggregation for AI | Variable |
| iaskspideriAsk.ai | AI search indexing | Yes |
| img2datasetHugging Face | Image dataset collection | No |
The bottom four rows are where good intentions run out. For bots with poor compliance records, robots.txt is a request rather than a control. If you genuinely need to stop them, that happens at the server or CDN with a 403, not in a text file. | ||
No crawler matches that. Try a company name, or clear the filter.
The four that matter most
If you only act on four lines of this page:
Decides whether ChatGPT can retrieve you. ChatGPT reaches roughly 900 million people a week. This is the highest-cost block on the list.
Opens ChatGPT →Serves Copilot and feeds a large share of ChatGPT’s citations. One token, two major platforms.
Opens Copilot and ChatGPT →Serves AI Overviews, AI Mode and Gemini’s grounded answers alongside classic search. You almost certainly already allow it, but check that a staging rule or a security plugin has not narrowed it.
Opens AI Overviews, AI Mode, Gemini →Your entry into the platform where SEO investment transfers most directly, and where results usually arrive fastest.
Opens Perplexity →Training vs search: the decision you actually have
Most businesses should allow everything. If your goal is to be found and recommended, there is no upside to withholding content from training: it is the raw material from which models learn your brand exists.
Publishers are the real exception, and the calculation is genuinely different. If your content is the product, and an assistant summarizing it means nobody visits, then training exclusion is a rational commercial position rather than a reflex.
Here is the part that gets missed. Those two decisions are separable.
The middle route is the one almost nobody knows exists. You keep every citation opportunity and remove your content from future training datasets. For a publisher weighing this, it is usually the right answer, and it is available today with a handful of lines.
One caveat worth being straight about: blocking training bots now does nothing about models already trained. That content is in. This is a forward-looking control, not an eraser.
Why robots.txt is not the whole story
This is the section most people need, and the reason accidental blocks are so common.
Your robots.txt can be perfect and your site can still be unreadable, because the request never reaches robots.txt in the first place.
The request is refused before it reaches robots.txt. Your file can be perfect and the site still unreadable.
Where blocks actually happen:
| Layer | Typical cause | Visible in robots.txt? |
|---|---|---|
| CDN bot protection | Cloudflare bot-fight mode, AI crawler blocking toggles, often on by default | No |
| WAF rules | Aggressive rate limiting or bot heuristics | No |
| Security plugins | WordPress plugins blocking unfamiliar agents out of the box | No |
| Server config | Old .htaccess or nginx rules nobody remembers adding | No |
| JavaScript rendering | Content that only exists after execution | Not a block, but the same effect |
That last row deserves emphasis. Most AI crawlers do not execute JavaScript. A client-side rendered site is not blocked, exactly. It is simply empty when read.
Server-side rendering is not a nice-to-have for AI visibility, it is the difference between having content and appearing to have none.
Several CDN providers added AI crawler blocking as a default or near-default setting during 2025 and 2026. A significant share of the accidental blocks we find trace back to a toggle nobody consciously flipped.
How to actually check
Ten minutes, in this order.
Read your robots.txt properly
Visit yourdomain.com/robots.txt and look for any Disallow: / under an AI token, plus any blanket User-agent: * rule that catches everything.
Check every subdomain separately. Rules do not inherit. shop.yourdomain.com and blog.yourdomain.com each need their own file.
Check your CDN and firewall
In Cloudflare, look at Security, then Bots, and specifically at any AI crawler control. Other providers have equivalents under bot management. This is where the surprises are.
Read your server logs
The only reliable evidence. Everything else is what you have configured; logs are what actually happened.
Search the last 30 days for these tokens: GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, bingbot, Googlebot.
You are looking for two things: whether they visited at all, and what status codes they received. A pattern of 403s is a block. Silence is worse, because it usually means they never got that far.
Test your rendering
Fetch a key page with JavaScript disabled, or use a plain HTTP request rather than a browser. If the page comes back empty, that is what most AI crawlers see.
If steps 1 and 2 look fine but the logs show nothing, the problem is upstream of both. That is the moment to bring in someone who reads logs for a living. Once crawlers can reach you, the next question is whether AI actually names you, which is what measuring AI visibility answers.
Should you block anything?
A balanced view, because the reflexive answers in both directions are wrong.
Reasons to allow everything:
- You want to be found, cited and recommended
- Training exposure is how models learn your brand exists at all
- The cost of accidental over-blocking is far higher than the cost of being trained on
Reasons a publisher might block training bots:
- Your content is the product and summarization substitutes for visiting
- You are pursuing or maintaining a licensing position
- You hold rights that make broad reuse genuinely problematic
Reasons almost nobody should block search bots:
- It removes you from AI answers entirely, with no compensating benefit
Reasons blocking may not work anyway:
- Several crawlers have poor compliance records
- Some data reaches models through third-party datasets you never see
- Content already used in training cannot be withdrawn
If you do block, block precisely. A blanket rule is how a business intending to protect its archive accidentally removes itself from ChatGPT.
Seven common mistakes
| The mistake | What it costs | The fix |
|---|---|---|
| Blocking GPTBot to “opt out of AI” | Nothing gained on visibility, and it does not affect search retrieval anyway | Understand the three jobs before writing any rule |
| Assuming robots.txt is the whole picture | Silent CDN block, invisible in every dashboard | Check CDN and firewall settings directly |
| Forgetting subdomains | Whole sections unreadable | One robots.txt per subdomain |
| Client-side rendering only | Crawlers see an empty page | Server-side render |
| Blocking by IP range | Unreliable, most vendors do not publish ranges, can block robots.txt itself | Match on user agent instead |
| Copying a robots.txt from a blog post | Stale tokens that match nothing | Verify tokens against vendor documentation |
| Never checking the logs | No idea whether any of it worked | Quarterly log review |
Structured data is the next layer down once crawlers can read you: see our guide to structured data for AI.
Frequently asked questions
What is the difference between GPTBot and OAI-SearchBot?
Does Google-Extended stop me appearing in AI Overviews?
How do I know if I am blocking AI crawlers?
Do AI crawlers respect robots.txt?
Should I block AI crawlers?
Can AI crawlers read JavaScript?
How often should I check this?
Does blocking a training bot remove my content from models already trained?
Not sure what is reaching your site?
Crawler access is the first thing we check on any audit, because it is binary. If the models cannot read you, nothing else you do matters.
We check your robots.txt across every subdomain, your CDN and firewall rules, and 30 days of server logs, then show you exactly which crawlers reached you and what they got back.