Reference
All 70 checks, one page each
Every check the scanner runs, what it is worth, and how much of the web currently fails it. Each one has its own page saying what it measures, why an agent cares, and the exact lines to paste.
The weights and the gates are in the specification; this is the index. The last column is measured rather than asserted: of 292 public sites we scanned on 2026-09-06, how many failed or warned on that check, out of the ones it applied to at all. Checks are ordered within each pillar by that count, so the ones most likely to be your problem come first.
Reachability
| Check | Weight | Sites failing |
|---|---|---|
| Re-fetching is cheap for a crawler | 4 | 162 of 292 |
| Content exists without running JavaScript | 14 | 79 of 292 |
| robots.txt is present and parseable | 4 | 59 of 292 |
| Declares a correct content type | 3 | 57 of 292 |
| Page responds to a plain HTTP request | 10 | 48 of 292 |
| Answer engines are allowed to crawl | 12 | 47 of 292 |
| No bot wall in front of the content | 12 | 45 of 292 |
| User-triggered fetchers are allowed | 8 | 44 of 292 |
| Retrieval providers are allowed | 5 | 36 of 292 |
| No directive suppressing AI use of this page | 6 | 24 of 292 |
| Responds fast enough for an agent budget | 5 | 18 of 292 |
| HTML is compressed on the wire | 3 | 13 of 292 |
| Autonomous agents are allowed | 8 | 12 of 292 |
| The status code matches what the page says | 5 | 9 of 244 |
| HTML payload stays inside an agent context budget | 5 | 7 of 292 |
| Short redirect chain | 4 | 4 of 292 |
| Classic search crawlers are allowed | 6 | 1 of 292 |
| Coding agents can read the documentation | 4 | 0 of 21 |
| Real agent user-agents are served normally | 8 | — |
Comprehension
| Check | Weight | Sites failing |
|---|---|---|
| Signal-to-markup ratio | 8 | 252 of 292 |
| Uses semantic elements, not div soup | 6 | 194 of 231 |
| Headings form a sane outline | 6 | 142 of 292 |
| Answers questions in a form that can be quoted | 8 | 129 of 191 |
| Has a meta description | 8 | 125 of 292 |
| Exactly one <h1> naming the page | 8 | 121 of 292 |
| Images carry alt text | 8 | 109 of 220 |
| Enough substance to be worth citing | 10 | 101 of 292 |
| Main content is marked as such | 6 | 87 of 292 |
| Has a descriptive <title> | 10 | 76 of 292 |
| Passages make sense on their own | 7 | 38 of 261 |
| Declares a document language | 5 | 31 of 292 |
| Provides a no-JavaScript fallback | 4 | 11 of 11 |
| Text arrived intact, not mangled | 5 | 5 of 292 |
| The page describes itself consistently | 6 | 3 of 191 |
| URL describes the content | 3 | 0 of 292 |
Structure
| Check | Weight | Sites failing |
|---|---|---|
| Declares a canonical URL | 6 | 239 of 292 |
| Ships structured data | 12 | 163 of 292 |
| Open Graph card is complete | 6 | 158 of 292 |
| Declares a link preview card | 3 | 116 of 292 |
| Structured data describes the right thing | 8 | 106 of 129 |
| Publishes a sitemap | 8 | 96 of 292 |
| Dates are machine-readable | 3 | 69 of 114 |
| Prices are machine-readable | 6 | 68 of 70 |
| Structured data agrees with the page | 7 | 37 of 129 |
| Declares its language alternates | 4 | 16 of 93 |
| Structured data parses cleanly | 8 | 10 of 129 |
| Tabular data is marked up as a table | 4 | 7 of 8 |
| Publishes a feed of its content | 4 | 5 of 58 |
| Declares its place in the site | 4 | — |
| Inline microdata annotations | 2 | 0 of 14 |
Actionability
| Check | Weight | Sites failing |
|---|---|---|
| Site search is machine-usable | 4 | 275 of 292 |
| Programmatic surface is discoverable | 6 | 267 of 292 |
| Declares what an agent can do here | 5 | 248 of 292 |
| Reachable by a human, discoverably | 5 | 231 of 292 |
| Navigation uses real links | 12 | 142 of 292 |
| Fields declare autocomplete tokens | 4 | 111 of 116 |
| Link text describes the destination | 8 | 83 of 226 |
| Links onward into the site | 6 | 73 of 292 |
| Controls are real buttons | 6 | 59 of 209 |
| Form fields are labelled | 8 | 50 of 120 |
| Form fields have stable names | 6 | 16 of 120 |
Contract
| Check | Weight | Sites failing |
|---|---|---|
| Declares AI usage preferences the standard way | 5 | 292 of 292 |
| Declares content licensing | 5 | 280 of 292 |
| Exposes an agent-callable manifest | 6 | 267 of 292 |
| Publishes an llms.txt | 14 | 228 of 292 |
| Says when the content changed | 5 | 224 of 292 |
| Identity is machine-verifiable | 5 | 223 of 292 |
| States an explicit AI policy | 6 | 212 of 292 |
| Publishes security.txt | 4 | 181 of 292 |
| llms.txt is actually useful | 6 | 7 of 64 |
Which check should I fix first?
The one at the top of your own report, not the one at the top of this page. A report orders findings by how many points each fix recovers on your page, which combines the weight with how badly you failed it and whether it applies to you at all. This index is ordered by how common a failure is across the web, which is a different question and a good way to find out that the thing you thought was exotic is the thing everybody gets wrong.
How is a check's weight decided?
By what it costs an agent, not by how hard it is to fix. Reaching the page at all is worth more than describing it well, because a page an agent cannot fetch scores nothing on everything else. Every weight is published in the specification and the rubric is versioned, so a score can be compared with the same score from a month ago.
Where does the failure count come from?
From 292 well-known public sites, each fetched once without a browser on 2026-09-06 and scored against this same rubric. The whole corpus is published as an open dataset under CC BY 4.0, method and limits included. No site in it is named.
Which of these are you failing? One scan, about two seconds, no account.