Everything on this site, and what each thing answers
62 destinations: 4 tools that read a page you own, 40 pages describing the individual rules, 12 describing the dimensions those rules are scored in, 5 guides for the fixes that recur, and the method they all rest on.
The tools are free and need no account. The reference pages are free too, which is the unusual half: the rules, their points and their pass conditions are published rather than described, so a score can be argued with instead of taken on trust.
| If you are asking | Go to | Destinations |
|---|---|---|
| I want a number for my own site | The tools themselves | 4 |
| I was told a rule failed | The published rule reference | 41 |
| I want to know where the number comes from | The twelve weighted dimensions | 12 |
| I want to check the method rather than trust it | The method and the evidence behind it | 1 |
| Something failed and I need to fix it | The guides | 5 |
The tools themselves
I want a number for my own site · 4 destinations
- GEO audit reportScores one homepage against the published checks and lists what failed with the evidence each check produced.Free. No account. No email. Nothing about the site is stored, and the result exists only in the URL you are on.
- /llms.txt studioDrafts an llms.txt from the site's own markup: the real page title, the homepage description, and up to 8 links per section, each described with that linked page's own meta description or H1 where it was read.No description in the file is written by the tool. It reads the homepage and up to 12 of the pages it links to, then says in the file how it was produced and what it could not read.
- GEO readiness badgeEmbeds a badge carrying the score from a real scan, as Markdown for a README or HTML for a footer. The number is read from the scan and cannot be typed in.No account, no email, and nothing to install.
- Weekly GEO reportEmails you when a site's score changes: the same checks run once a week, naming the checks that newly fail, with the evidence the failing ones saw.No account. The email address is the whole subscription - confirming it starts the report and the unsubscribe link ends it.
The published rule reference
I was told a rule failed · 41 destinations
- All 40 checksEvery rule the scanner runs, grouped by the dimension it belongs to, each with its exact pass condition and its point value.
- robots.txt is servedA request for /robots.txt returns HTTP 200 with a non-empty body.
- No AI crawler is blockedParsing robots.txt into user-agent groups, none of gptbot, claudebot, perplexitybot, oai-searchbot or google-extended is disallowed from /, AND the homepage served a crawler-shaped request with HTTP 200.
- robots.txt declares a sitemaprobots.txt contains a Sitemap: line.
- Homepage returns HTTP 200The homepage returns HTTP 200 to a non-browser request that identifies itself honestly.
- Homepage is marked noindexThe homepage has no meta robots directive containing noindex.
- sitemap.xml is present and well-formed/sitemap.xml returns a body containing <urlset> or <sitemapindex>.
- robots.txt states a content policyrobots.txt contains a Content-Signal directive with a value on the same line, for example `Content-Signal: search=yes, ai-input=yes, ai-train=no`.
- Valid JSON-LD is presentAt least one <script type="application/ld+json"> block parses as JSON and declares at least one @-keyword (@context, @type, @graph or @id).
- An entity node is declaredParsed JSON-LD contains an @type of Organization, WebSite, Person or LocalBusiness.
- Content-type schema is presentParsed JSON-LD contains an @type of FAQPage, Article, BlogPosting, HowTo, Product, SoftwareApplication or BreadcrumbList.
- Substantial body textVisible text, after removing script, style and comment content, contains at least 800 word tokens.
- Content is divided into sectionsThe page contains at least 4 H2 elements.
- Lists and tables are presentThe page contains both a <ul>/<ol> and a <table>.
- Quantified claims are presentBody text contains at least 5 numeric claims: percentages, currency amounts, multipliers of the form 3.2x, thousands-separated figures, or raw numbers of five digits or more.
- Direct quotations are presentThe page contains at least one <blockquote> element.
- External authoritative sources are citedThe page links to at least 2 external hosts on the authority list: arxiv.org, doi.org, nature.com, science.org, acm.org, ieee.org, springer.com, sciencedirect.com, any .gov or .edu host, wikipedia.org, github.com, developer.mozilla.org, ...
- Canonical URL is declaredThe page declares a rel=canonical link.
- Headings are phrased as real questionsAt least 3 H2 or H3 headings either end with a question mark or begin with how, what, why, which, when, who, where, is, are, do, does, can, should or will.
- FAQPage schema is presentParsed JSON-LD contains a FAQPage node.
- Headings are followed by a direct answerOf the first 30 heading-then-paragraph pairs (H2/H3 optionally followed by wrapper elements then a <p>), at least 60% have an opening paragraph of 80 words or fewer.
- Served over HTTPSThe homepage was retrieved over https://, rather than only over plain http://.
- About and contact paths are linkedHrefs on the page include both an about-style path (about, company, team, who-we-are) and a contact-style path (contact, support) or a mailto: link.
- Authorship is attributedAuthorship is signalled by any of: a Person node in parsed JSON-LD, an author property with a non-empty value in parsed JSON-LD, rel="author", or a visible byline matching "by Firstname Lastname".
- Entity is cross-referencedParsed JSON-LD contains a sameAs property with a non-empty value.
- Exactly one H1The page contains exactly one non-empty <h1>.
- Semantic landmarks are usedAt least 3 of main, article, section, header, nav and footer appear as elements.
- Main content is in the HTMLThe raw server response contains at least 100 words of visible text without executing JavaScript.
- Title length is in range<title> is between 15 and 65 characters long.
- Meta description length is in rangeThe meta description is between 50 and 160 characters long.
- Open Graph tags are presentAt least one meta property beginning with og: is present.
- Document language is declaredThe <html> element carries a lang attribute.
- /llms.txt is present and structured/llms.txt returns more than 20 bytes and contains an H1, a > summary line and at least 3 markdown links.
- A crawler policy is publishedrobots.txt is readable, so a crawler policy is published.
- A markdown alternate is declaredThe document head contains a <link> with rel="alternate" and type="text/markdown", in either attribute order.
- A machine-readable date is present and recentA machine-readable date is found in dateModified, datePublished or article:modified_time.
- A current year appears in the contentThe current calendar year appears in the visible body text.
- Alternate language versions are declaredAt least 2 distinct hreflang language codes are declared.
- A region-specific locale is declaredA language tag carrying a region subtag appears in <html lang>, in og:locale or in an hreflang attribute - for example en-GB, en_GB or de-AT.
- Viewport meta tag is presentA meta viewport tag is present.
- HTML payload is reasonableThe uncompressed HTML response is under 500 KB.
The twelve weighted dimensions
I want to know where the number comes from · 12 destinations
- AI Crawler Access — 16% of the scoreIf a crawler cannot read the page, nothing else on this list can matter. This is the only dimension where a failure makes the rest of the score moot.
- Machine Readability — 12% of the scoreJSON-LD is how a model binds your brand name, domain and product into one entity instead of inferring three unrelated strings.
- Content Depth — 11% of the scoreGenerative engines select sources that answer a question thoroughly. Thin pages are rarely retrievable regardless of how well they are marked up.
- Citability & Evidence — 11% of the scoreStatistics, quotations and cited sources are the interventions with the largest measured effect in the published GEO research.
- Answer Readiness — 10% of the scoreAnswers are extracted as spans, not pages. A question-shaped heading followed by an immediate answer is the easiest thing for a retrieval system to lift intact.
- Trust & Authority — 10% of the scoreE-E-A-T signals decide whether a model treats a claim as safe to repeat rather than something it should hedge.
- Semantic Structure — 8% of the scoreHeading hierarchy and landmark elements are how a parser locates section boundaries at all.
- Metadata & Discoverability — 7% of the scoreTitle, description and canonical control what a search or answer surface can show about the page.
- AI Context Files — 5% of the scoreA cheap and optional signal. Google has stated it does not use llms.txt in Search, and crawler support is inconsistent.
- Freshness — 5% of the scoreDated content is deprioritised in generated answers, and an undated page gives an engine nothing to reason about.
- International Readiness — 3% of the scoreLanguage and region markup decides which language market a page can be retrieved in at all.
- Delivery & Mobile — 2% of the scoreA slow or non-mobile-readable page is dropped before any content analysis happens.
The method and the evidence behind it
I want to check the method rather than trust it · 1 destination
The guides
Something failed and I need to fix it · 5 destinations
- How to check whether AI engines mention your brand
- How to generate and deploy an llms.txt file
- Optimizing headings for direct AI citation
- Implementing Schema.org JSON-LD for entity disambiguation
- Configuring robots.txt and WAF for AI crawlers
- 16 of the 40 rule pages name the guide that covers their fix. The rest show the fix in full on the rule page itself, because no guide covers them and pointing at one that half applies would be worse than saying nothing.
What this index leaves out
The scanner reads one page. It does not query ChatGPT, Perplexity, Gemini or any other engine, so there is no tool here for asking what an engine answers - the self-check guide is the honest version of that, done by hand.
The scan result page is not listed either, because it is generated per request from the domain you type rather than published. Every rule above links into that report, so the fastest way to see which of these 40 rules a page currently fails is to run one against your own domain. The scoring weights across the 12 dimensions add up to 100.
Questions about these tools and references
Do I need an account for any of this?
No. The 4 tools read a page you own and report what they found, and the 40 rule pages, the 12 dimension pages and the method are published documents. The one thing that takes an address is the weekly report, and confirming that address is the whole subscription.
Where do I start if I only have five minutes?
Run the audit report on your own domain. It returns a score out of 100, the twelve dimension scores behind it and the list of what failed, and every failed rule links to its own page here. Reading this index first is for when you know which question you have.
Why do only 16 of the rule pages link to a guide?
The 5 guides cover the fixes that recur across sites: crawler access, schema, headings, llms.txt and how to check citations by hand. The other 24 rules have no guide, so their pages carry the full fix rather than a link to one that only half applies.
Is there an API?
No. The scanner is a public endpoint behind rate limiting rather than a documented API, and this page does not pretend otherwise. What is published instead is the method itself - every rule, its points and its pass condition - so the score can be recomputed rather than trusted.
Evidence and sources
Adding source citations produced the largest measured visibility gain for low-ranking sites, at +115%, ahead of the addition of expert quotations at +41% and statistics at +30-40%, across the strategies tested on generative engines. — Generative Engine Optimization, KDD 2024
The weighting across the 12 dimensions follows that measurement rather than taste, and the rule set is published in full so a score can be recomputed by hand rather than taken on trust.
Primary sources
- Generative Engine Optimization (KDD 2024) — the citation and quotation figures the citability weight follows
- What Generative Search Engines Like — which page characteristics are surfaced in generated answers
- What Gets Cited: Competitive GEO — earned media is favoured over brand-owned content
- The rule set and its weights — every check, its weight and its pass condition