AXRAY

Reference

Which AI crawlers should you actually block?

Every crawler below is one this scanner probes with, so the list is the product rather than a copy of somebody else’s. The last column is measured: it is how many sites in a sample of 293 public websites already block that crawler in robots.txt.

Türkçe

The distinction that decides everything

Answer engines

These read you so an assistant can cite you and link you. Blocking one removes you from that assistant’s answers entirely - not "less visible", absent.

User-triggered fetchers

These fetch a page because a person asked an assistant to open it, right now, with the person waiting. Blocking one means the link somebody pasted fails in front of them.

Training crawlers

These collect content to train models. Blocking them is a legitimate business decision and it costs you no traffic, which is exactly what makes it the safe one to take.

The full list

CrawlerOperatorWhat it is forAlready blocked by
OAI-SearchBot OpenAI Answer engines 20 of 293
Claude-SearchBot Anthropic Answer engines 26 of 293
PerplexityBot Perplexity Answer engines 36 of 293
Amazonbot Amazon Answer engines 35 of 293
DuckAssistBot DuckDuckGo Answer engines 20 of 293
ChatGPT-User OpenAI User-triggered fetchers 27 of 293
Claude-User Anthropic User-triggered fetchers 27 of 293
Perplexity-User Perplexity User-triggered fetchers 23 of 293
Googlebot Google Classic search crawlers 0 of 293
Bingbot Microsoft Classic search crawlers 1 of 293
GPTBot OpenAI Training crawlers 41 of 293
ClaudeBot Anthropic Training crawlers 43 of 293
Google-Extended Google Training crawlers 37 of 293
Applebot-Extended Apple Training crawlers 37 of 293
meta-externalagent Meta Training crawlers 34 of 293
CCBot Common Crawl Training crawlers 55 of 293

Blocking figures measured on 2026-09-03 across 293 public websites. No site is named.

Each one, and what blocking it costs

OAI-SearchBot

OpenAI · Answer engines

Blocking it means: You will not appear as a source in ChatGPT search results.

User agent: Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)

20 of 293 · 53 sites have no robots.txt rule either way.

Claude-SearchBot

Anthropic · Answer engines

Blocking it means: You will not be cited in Claude web search answers.

User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; +Claude-SearchBot@anthropic.com)

26 of 293 · 53 sites have no robots.txt rule either way.

PerplexityBot

Perplexity · Answer engines

Blocking it means: You lose Perplexity citations, one of the highest-converting AI referral sources.

User agent: Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

36 of 293 · 53 sites have no robots.txt rule either way.

Amazonbot

Amazon · Answer engines

Blocking it means: Excluded from Alexa and Rufus style shopping answers.

User agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)

35 of 293 · 53 sites have no robots.txt rule either way.

DuckAssistBot

DuckDuckGo · Answer engines

Blocking it means: Excluded from DuckDuckGo AI assist answers.

User agent: Mozilla/5.0 (compatible; DuckAssistBot/1.0; +https://duckduckgo.com/duckassistbot/)

20 of 293 · 53 sites have no robots.txt rule either way.

ChatGPT-User

OpenAI · User-triggered fetchers

Blocking it means: When a user asks ChatGPT to open your link, it fails in front of them.

User agent: Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)

27 of 293 · 53 sites have no robots.txt rule either way.

Claude-User

Anthropic · User-triggered fetchers

Blocking it means: Claude cannot open your page when a user pastes the link.

User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com)

27 of 293 · 53 sites have no robots.txt rule either way.

Perplexity-User

Perplexity · User-triggered fetchers

Blocking it means: Live user-requested fetches from Perplexity fail.

User agent: Mozilla/5.0 (compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

23 of 293 · 53 sites have no robots.txt rule either way.

Googlebot

Google · Classic search crawlers

Blocking it means: You are invisible to Google Search and to AI Overviews built on it.

User agent: Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

0 of 293 · 53 sites have no robots.txt rule either way.

Bingbot

Microsoft · Classic search crawlers

Blocking it means: You are invisible to Bing and to Copilot answers grounded on it.

User agent: Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)

1 of 293 · 53 sites have no robots.txt rule either way.

GPTBot

OpenAI · Training crawlers

Blocking it means: Your content is excluded from OpenAI model training.

User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)

41 of 293 · 53 sites have no robots.txt rule either way.

ClaudeBot

Anthropic · Training crawlers

Blocking it means: Your content is excluded from Anthropic model training.

User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)

43 of 293 · 53 sites have no robots.txt rule either way.

Google-Extended

Google · Training crawlers

Blocking it means: Excluded from Gemini training and grounding.

User agent: Mozilla/5.0 (compatible; Google-Extended/1.0)

37 of 293 · 53 sites have no robots.txt rule either way.

Applebot-Extended

Apple · Training crawlers

Blocking it means: Excluded from Apple Intelligence training.

User agent: Mozilla/5.0 (compatible; Applebot-Extended/1.0)

37 of 293 · 53 sites have no robots.txt rule either way.

meta-externalagent

Meta · Training crawlers

Blocking it means: Excluded from Meta AI training.

User agent: meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)

34 of 293 · 53 sites have no robots.txt rule either way.

CCBot

Common Crawl · Training crawlers

Blocking it means: Excluded from Common Crawl, the base corpus behind most open models.

User agent: CCBot/2.0 (https://commoncrawl.org/faq/)

55 of 293 · 53 sites have no robots.txt rule either way.

A robots.txt that separates the two decisions

This is the shape worth copying: refuse training, keep every crawler that could send you a visitor. Change the last group if you want to be trained on - that is the only line in it that is a matter of opinion.

# Everything that could send you a visitor. Keep these.
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Amazonbot
User-agent: DuckAssistBot
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: Googlebot
User-agent: Bingbot
Allow: /

# Training crawlers. This is the only group that is a matter of opinion.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /

Sitemap: https://axray.online/sitemap.xml

One rule that people get wrong: a crawler matches the most specific user-agent group that names it, and nothing else. A crawler named in its own group ignores your User-agent: * block entirely, including the Disallow lines you wanted to keep.

Questions people actually ask

Does blocking GPTBot stop ChatGPT from citing my site?

No. GPTBot collects content for training. The crawler that puts you into ChatGPT answers is OAI-SearchBot, and the one that opens a link when somebody asks ChatGPT to read it is ChatGPT-User. Blocking all three is a common accident, because most published robots.txt advice lists them together as "the OpenAI bots".

Is robots.txt legally binding on AI crawlers?

It is a request, not a fence, and it has never been anything else - the same was true of search crawlers for twenty-five years. The major operators listed here document the tokens they honour. If you need enforcement rather than a request, that is a firewall rule or a licence agreement, not a text file.

What happens if I have no robots.txt at all?

Everything is allowed, which is the default and is usually fine. In the sample behind this page, 53 of 293 sites had no rule either way for the main training crawlers. The risk of an absent file is not that you are crawled; it is that you have made no decision and cannot tell anyone what your decision was.

Will blocking AI crawlers protect my content from being used?

Only by the operators that honour robots.txt, which is the well-behaved subset. It does nothing about content that reaches a model through a third-party copy, a screenshot, or a user pasting your text into a chat window. Blocking is worth doing as a stated position; it is not a technical guarantee and nobody should sell it to you as one.

See which of these your own robots.txt currently blocks, and what each one costs you - in one scan, with no account.