AXRAY
Sign inCreate account

Classic search

Should you block Googlebot?

Googlebot is operated by Google. Traditional search indexing, increasingly the substrate AI answers are built on.

All 54 AI crawlers

What Googlebot is

Google runs it as a classic search crawler, which means it does ordinary search indexing, of the kind that has existed for decades. That matters more than it used to rather than less: several assistants answer by searching a conventional index first and reading the top results.

What blocking it costs you: You are invisible to Google Search and to AI Overviews built on it. Ordinary search indexing stops, and increasingly that index is the substrate AI answers are built from as well.

The user agent, and the token a rule matches on

These are two different strings and confusing them is why a rule that looks right does nothing. A robots.txt rule matches on the product token, case-insensitively, and ignores the rest of the user-agent header entirely.

robots.txt tokenGooglebot
OperatorGoogle
CategoryClassic search
Full user agentMozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

Allowing Googlebot

Paste this into robots.txt at the root of your domain.

User-agent: Googlebot
Allow: /

Sitemap: https://axray.online/sitemap.xml

An empty Disallow: and Allow: / mean the same thing to a crawler. What does not mean the same thing is having no rule at all: absence is permission by default, but it is permission nobody wrote down, and it survives only until somebody adds a blanket User-agent: * block without thinking about this crawler.

Blocking Googlebot

User-agent: Googlebot
Disallow: /

Before you paste that: you are invisible to Google Search and to AI Overviews built on it. That is the cost, and it is paid silently — no error appears in your logs, no report tells you, and the traffic you lose was traffic you never saw arrive.

How the web actually treats Googlebot

Measured on 2026-09-06, by reading the robots.txt of 292 public websites and resolving each one against this crawler specifically.

VerdictSitesShare
Blocked00%
Explicitly allowed23882%
No rule either way5418%

No site is named, here or anywhere else on this domain. The count is the useful part; a league table of businesses that never asked to be measured is not.

The row worth reading twice is the last one. 54 of 292 sites have written no rule about Googlebot at all, which means their position on it is an accident rather than a decision — whatever their User-agent: * block happens to say.

What Googlebot is not

Google runs 6 other crawlers with different jobs, and this is where the expensive mistake happens. Somebody means “do not train on my content”, writes one rule against the operator’s name as they remember it, and blocks the crawler that would have sent them a customer instead.

CrawlerWhat it does instead
Google-NotebookLM Fetches your page because a human asked an assistant to open it right now. Blocking it breaks a live request.
GoogleAgent-URLContext Fetches your page because a human asked an assistant to open it right now. Blocking it breaks a live request.
GoogleAgent-Mariner Browses, compares and buys on a person’s behalf. Blocking it removes you from the shortlist before anyone sees it.
Gemini-Deep-Research Browses, compares and buys on a person’s behalf. Blocking it removes you from the shortlist before anyone sees it.
Google-Gemini-CLI Reads documentation while a developer is building against you. Blocking it is how your API gets used wrong.
Google-Extended Collects content for model training. Blocking it is a legitimate business choice with no traffic cost.

Rules are matched per token. Blocking one of these says nothing about the others, which is the point: you can refuse training and keep every crawler that puts your link in front of a person.

Checking what your site does right now

Reading your own robots.txt is not the same as knowing what a crawler concludes from it. Precedence between a User-agent: * group and a named one, the longest-match rule between overlapping paths, and a token you spelled slightly wrong all produce a file that looks correct and behaves otherwise.

npx axray-cli your-site.com

The scan resolves your file against all 54 crawlers on this list individually and reports which ones are allowed, blocked, or covered by nothing. It is free, needs no account, and installs nothing.

Questions people ask about Googlebot

What is Googlebot?

Googlebot is a classic search crawler operated by Google: it does ordinary search indexing, of the kind that has existed for decades. That matters more than it used to rather than less: several assistants answer by searching a conventional index first and reading the top results.

Should I block Googlebot in robots.txt?

Only if you are willing to pay what it costs. You are invisible to Google Search and to AI Overviews built on it. Ordinary search indexing stops, and increasingly that index is the substrate AI answers are built from as well.

What is the Googlebot user agent string?

Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) — and the robots.txt product token, which is what a rule actually matches on, is Googlebot.

How do I check whether my site is blocking Googlebot right now?

Scan your site with AXRAY. It resolves your robots.txt against 54 named AI crawlers individually rather than reporting one verdict for all of them, so it will tell you whether this specific crawler is allowed, blocked, or covered by no rule at all. It is free and needs no account.

Other classic search crawlers

The full reference: all 54 AI crawlers, and who blocks each one.

See which of these 54 crawlers your own site is blocking, and what each one costs you. One scan, no account.