What is AI2Bot?
AI2Bot is a training crawler operated by Ai2: it collects text to train a model on. It sends nobody to your site, now or later, and it is the only category on this list where that is true.
Training crawler
AI2Bot is operated by Ai2. Collects content for model training. Blocking it is a legitimate business choice with no traffic cost.
Ai2 runs it as a training crawler, which means it collects text to train a model on. It sends nobody to your site, now or later, and it is the only category on this list where that is true.
What blocking it costs you: Excluded from the open training corpora published by the Allen Institute, which are what most academic and open-weight models are built from. Nothing, in traffic terms. This is the one group where blocking has a real argument behind it and no visitor cost.
These are two different strings and confusing them is why a rule that looks right does nothing.
A robots.txt rule matches on the product token, case-insensitively, and
ignores the rest of the user-agent header entirely.
| robots.txt token | AI2Bot |
|---|---|
| Operator | Ai2 |
| Category | Training crawler |
| Full user agent | Not published by Ai2. Match on the token above. |
Where an operator has not published the exact string, this page says so rather than inventing a plausible one. A user agent you can grep your logs for is worth nothing if it is a guess.
Paste this into robots.txt at the root of your domain.
User-agent: AI2Bot
Allow: /
Sitemap: https://axray.online/sitemap.xml
An empty Disallow: and Allow: / mean the same thing to a crawler. What does
not mean the same thing is having no rule at all: absence is permission by default, but it is
permission nobody wrote down, and it survives only until somebody adds a blanket
User-agent: * block without thinking about this crawler.
User-agent: AI2Bot
Disallow: /
This one is a fair decision to make. A training crawler sends you no visitors, so blocking it costs you nothing in traffic and is a position you are entitled to take. Be clear about what it achieves: robots.txt is honoured by operators who choose to honour it, and does nothing about your text reaching a model through a third-party copy, a screenshot, or somebody pasting it into a chat window.
Measured on 2026-09-06, by reading the robots.txt of
292 public websites and resolving each one against this crawler specifically.
| Verdict | Sites | Share |
|---|---|---|
| Blocked | 18 | 6% |
| Explicitly allowed | 220 | 75% |
| No rule either way | 54 | 18% |
No site is named, here or anywhere else on this domain. The count is the useful part; a league table of businesses that never asked to be measured is not.
The row worth reading twice is the last one. 54 of 292 sites have written no rule about
AI2Bot at all, which means their position on it is an accident rather than a decision — whatever
their User-agent: * block happens to say.
Reading your own robots.txt is not the same as knowing what a crawler concludes from it.
Precedence between a User-agent: * group and a named one, the longest-match rule between
overlapping paths, and a token you spelled slightly wrong all produce a file that looks correct and
behaves otherwise.
npx axray-cli your-site.com
The scan resolves your file against all 54 crawlers on this list individually and reports which ones are allowed, blocked, or covered by nothing. It is free, needs no account, and installs nothing.
AI2Bot is a training crawler operated by Ai2: it collects text to train a model on. It sends nobody to your site, now or later, and it is the only category on this list where that is true.
Blocking AI2Bot costs you no visitors, because it collects content for model training rather than sending anyone to your site. That makes it the one group where the decision is genuinely yours to make on principle. What it does not do is keep your content out of models — it applies only to operators who read robots.txt, and does nothing about a copy reached through somebody else.
Ai2 does not publish a full user-agent string for this crawler. The robots.txt product token is AI2Bot, and that is what a rule matches on; matching is case-insensitive.
Scan your site with AXRAY. It resolves your robots.txt against 54 named AI crawlers individually rather than reporting one verdict for all of them, so it will tell you whether this specific crawler is allowed, blocked, or covered by no rule at all. It is free and needs no account.
GPTBot — OpenAIClaudeBot — AnthropicGoogle-Extended — GoogleApplebot-Extended — Applemeta-externalagent — MetaCCBot — Common CrawlBytespider — ByteDanceDeepSeekBot — DeepSeekThe full reference: all 54 AI crawlers, and who blocks each one.
See which of these 54 crawlers your own site is blocking, and what each one costs you. One scan, no account.