Why this dimension exists
If a crawler cannot read the page, nothing else on this list can matter. This is the only dimension where a failure makes the rest of the score moot.
Why it carries 16%
Highest weight of the twelve. A page that is blocked or unreadable has no path to being cited, so it dominates the total.
The rules it is scored by
Each rule is a single question the scanner asks about a page. Passing one adds its share of this dimension; failing it takes that share away. Every rule links to what it looks for and what evidence passes it.
- A request for /robots.txt returns HTTP 200 with a non-empty body. A 200 with an empty or whitespace-only body counts as not served, because it publishes no policy at all.1.9%
- Parsing robots.txt into user-agent groups, none of gptbot, claudebot, perplexitybot, oai-searchbot or google-extended is disallowed from /, AND the homepage served a crawler-shaped request with HTTP 200. An exact agent group takes precedence over the * group, the longest matching pattern wins and Allow wins ties, an empty Disallow value matches nothing, and * is treated as a wildcard. When the homepage refused the request, the verdict follows a second probe of the same URL made with a browser-shaped User-Agent: if that one returns 200 the refusal is aimed at identified crawlers, and the check fails under robots-ai-blocked; if it is refused as well, the refusal is about the scanning address rather than the site, and this check scores 2 of 6 as unverified.5.6%
- robots.txt contains a Sitemap: line.0.9%
- The homepage returns HTTP 200 to a non-browser request that identifies itself honestly.1.9%
- The homepage has no meta robots directive containing noindex.1.9%
- /sitemap.xml returns a body containing <urlset> or <sitemapindex>.2.8%
- robots.txt contains a Content-Signal directive with a value on the same line, for example `Content-Signal: search=yes, ai-input=yes, ai-train=no`. The directive is matched case-insensitively at the start of any line.0.9%
What this dimension cannot tell you
A single-URL scan reads one page. It cannot see the rest of the site, it cannot see how a model behaves over time, and it cannot see whether the page is actually cited for the queries that matter. Those are the parts of the picture a score of this kind is blind to, and they are listed in full on the methodology page rather than implied here.