Reference
Which AI crawlers should you actually block?
Every crawler below is one this scanner probes with, so the list is the product rather than a copy of somebody else’s. The last column is measured: it is how many sites in a sample of 293 public websites already block that crawler in robots.txt.
The distinction that decides everything
Answer engines
These read you so an assistant can cite you and link you. Blocking one removes you from that assistant’s answers entirely - not "less visible", absent.
User-triggered fetchers
These fetch a page because a person asked an assistant to open it, right now, with the person waiting. Blocking one means the link somebody pasted fails in front of them.
Classic search crawlers
Ordinary search indexing, which increasingly is also the substrate AI answers are built on. Almost nobody means to block these.
Training crawlers
These collect content to train models. Blocking them is a legitimate business decision and it costs you no traffic, which is exactly what makes it the safe one to take.
The full list
| Crawler | Operator | What it is for | Already blocked by |
|---|---|---|---|
OAI-SearchBot |
OpenAI | Answer engines | 20 of 293 |
Claude-SearchBot |
Anthropic | Answer engines | 26 of 293 |
PerplexityBot |
Perplexity | Answer engines | 36 of 293 |
Amazonbot |
Amazon | Answer engines | 35 of 293 |
DuckAssistBot |
DuckDuckGo | Answer engines | 20 of 293 |
ChatGPT-User |
OpenAI | User-triggered fetchers | 27 of 293 |
Claude-User |
Anthropic | User-triggered fetchers | 27 of 293 |
Perplexity-User |
Perplexity | User-triggered fetchers | 23 of 293 |
Googlebot |
Classic search crawlers | 0 of 293 | |
Bingbot |
Microsoft | Classic search crawlers | 1 of 293 |
GPTBot |
OpenAI | Training crawlers | 41 of 293 |
ClaudeBot |
Anthropic | Training crawlers | 43 of 293 |
Google-Extended |
Training crawlers | 37 of 293 | |
Applebot-Extended |
Apple | Training crawlers | 37 of 293 |
meta-externalagent |
Meta | Training crawlers | 34 of 293 |
CCBot |
Common Crawl | Training crawlers | 55 of 293 |
Blocking figures measured on 2026-09-03 across 293 public websites. No site is named.
Each one, and what blocking it costs
OAI-SearchBot
OpenAI · Answer engines
Blocking it means: You will not appear as a source in ChatGPT search results.
User agent: Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)
20 of 293 · 53 sites have no robots.txt rule either way.
Claude-SearchBot
Anthropic · Answer engines
Blocking it means: You will not be cited in Claude web search answers.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; +Claude-SearchBot@anthropic.com)
26 of 293 · 53 sites have no robots.txt rule either way.
PerplexityBot
Perplexity · Answer engines
Blocking it means: You lose Perplexity citations, one of the highest-converting AI referral sources.
User agent: Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
36 of 293 · 53 sites have no robots.txt rule either way.
Amazonbot
Amazon · Answer engines
Blocking it means: Excluded from Alexa and Rufus style shopping answers.
User agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)
35 of 293 · 53 sites have no robots.txt rule either way.
DuckAssistBot
DuckDuckGo · Answer engines
Blocking it means: Excluded from DuckDuckGo AI assist answers.
User agent: Mozilla/5.0 (compatible; DuckAssistBot/1.0; +https://duckduckgo.com/duckassistbot/)
20 of 293 · 53 sites have no robots.txt rule either way.
ChatGPT-User
OpenAI · User-triggered fetchers
Blocking it means: When a user asks ChatGPT to open your link, it fails in front of them.
User agent: Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)
27 of 293 · 53 sites have no robots.txt rule either way.
Claude-User
Anthropic · User-triggered fetchers
Blocking it means: Claude cannot open your page when a user pastes the link.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com)
27 of 293 · 53 sites have no robots.txt rule either way.
Perplexity-User
Perplexity · User-triggered fetchers
Blocking it means: Live user-requested fetches from Perplexity fail.
User agent: Mozilla/5.0 (compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
23 of 293 · 53 sites have no robots.txt rule either way.
Googlebot
Google · Classic search crawlers
Blocking it means: You are invisible to Google Search and to AI Overviews built on it.
User agent: Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
0 of 293 · 53 sites have no robots.txt rule either way.
Bingbot
Microsoft · Classic search crawlers
Blocking it means: You are invisible to Bing and to Copilot answers grounded on it.
User agent: Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)
1 of 293 · 53 sites have no robots.txt rule either way.
GPTBot
OpenAI · Training crawlers
Blocking it means: Your content is excluded from OpenAI model training.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)
41 of 293 · 53 sites have no robots.txt rule either way.
ClaudeBot
Anthropic · Training crawlers
Blocking it means: Your content is excluded from Anthropic model training.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)
43 of 293 · 53 sites have no robots.txt rule either way.
Google-Extended
Google · Training crawlers
Blocking it means: Excluded from Gemini training and grounding.
User agent: Mozilla/5.0 (compatible; Google-Extended/1.0)
37 of 293 · 53 sites have no robots.txt rule either way.
Applebot-Extended
Apple · Training crawlers
Blocking it means: Excluded from Apple Intelligence training.
User agent: Mozilla/5.0 (compatible; Applebot-Extended/1.0)
37 of 293 · 53 sites have no robots.txt rule either way.
meta-externalagent
Meta · Training crawlers
Blocking it means: Excluded from Meta AI training.
User agent: meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
34 of 293 · 53 sites have no robots.txt rule either way.
CCBot
Common Crawl · Training crawlers
Blocking it means: Excluded from Common Crawl, the base corpus behind most open models.
User agent: CCBot/2.0 (https://commoncrawl.org/faq/)
55 of 293 · 53 sites have no robots.txt rule either way.
A robots.txt that separates the two decisions
This is the shape worth copying: refuse training, keep every crawler that could send you a visitor. Change the last group if you want to be trained on - that is the only line in it that is a matter of opinion.
# Everything that could send you a visitor. Keep these.
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Amazonbot
User-agent: DuckAssistBot
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: Googlebot
User-agent: Bingbot
Allow: /
# Training crawlers. This is the only group that is a matter of opinion.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /
Sitemap: https://axray.online/sitemap.xml
One rule that people get wrong: a crawler matches the most specific user-agent group that names it, and nothing else. A crawler named in its own group ignores your User-agent: * block entirely, including the Disallow lines you wanted to keep.
Questions people actually ask
Does blocking GPTBot stop ChatGPT from citing my site?
No. GPTBot collects content for training. The crawler that puts you into ChatGPT answers is OAI-SearchBot, and the one that opens a link when somebody asks ChatGPT to read it is ChatGPT-User. Blocking all three is a common accident, because most published robots.txt advice lists them together as "the OpenAI bots".
Is robots.txt legally binding on AI crawlers?
It is a request, not a fence, and it has never been anything else - the same was true of search crawlers for twenty-five years. The major operators listed here document the tokens they honour. If you need enforcement rather than a request, that is a firewall rule or a licence agreement, not a text file.
What happens if I have no robots.txt at all?
Everything is allowed, which is the default and is usually fine. In the sample behind this page, 53 of 293 sites had no rule either way for the main training crawlers. The risk of an absent file is not that you are crawled; it is that you have made no decision and cannot tell anyone what your decision was.
Will blocking AI crawlers protect my content from being used?
Only by the operators that honour robots.txt, which is the well-behaved subset. It does nothing about content that reaches a model through a third-party copy, a screenshot, or a user pasting your text into a chat window. Blocking is worth doing as a stated position; it is not a technical guarantee and nobody should sell it to you as one.
See which of these your own robots.txt currently blocks, and what each one costs you - in one scan, with no account.