The 18 AI Visibility Checks Behind Your AI Readiness Score
Published · Updated · Engine 1.2.0 · 9 min read
The free Aura scan runs 18 deterministic checks in five categories and turns them into a 0-100 AI Readiness Score. This article lists every check, what it measures, how many points it carries, what kind of rule it is, and how to pass it. It also says plainly what the score does not tell you.
What the score is, and what it is not
The score is Aura's technical checklist: reproducible, site-side checks of whether AI systems can fetch, read and attribute a page. It is not a prediction of how often ChatGPT, Claude, Perplexity or Gemini will cite the site. No public evidence establishes a causal link between any single check below and citation frequency, and we do not claim one. Measuring actual mentions and citations takes repeated sampling of real answers, reported with sample sizes and confidence intervals; that is a separate product (Monitor), and even there the result is a rate with an interval, not a promise.
Two statements from Google are worth knowing, and worth not over-generalising: Google says its Search does not use llms.txt and that publishing one neither helps nor harms Search visibility, and that no special schema is required for its AI features (Google's AI optimization guide, AI features guidance). Those statements are Google's. Other providers have said less, and we do not assume they behave the same way.
Three kinds of check
Every finding on a report is labelled with one of three kinds, so you can tell what a failing check actually means:
| Kind | Meaning | Checks |
|---|---|---|
| Access blocker | Without it AI systems cannot fetch or trust the page at all. | site-reachable, ai-bots, https, status, speed |
| Best practice | A widely recommended best practice that makes a page easier to read, summarise and attribute. Not a guarantee of anything. | sitemap, jsonld, org-schema, rich-schema, og, canonical, title, meta-desc, h1, headings, quotable, depth |
| Optional convention | A proposed or optional convention that some tools honour and some, by their own statement, do not use. | llms-txt, feed |
Checks that cannot apply to a page are marked N/A and left out of that page's denominator: a marketing homepage with no FAQ is not asked for FAQ markup, and a site with no blog is not asked for a feed. The score should never push you to add markup that does not describe real content.
Scoring overview
| Category | # Checks | Max points | What it tests |
|---|---|---|---|
| AI Crawler Access | 1 | 30 | Are AI crawlers allowed to fetch the site? |
| Machine Discovery | 3 | 15 | Are the conventional discovery files present? |
| Structured Data | 5 | 20 | Does valid markup describe what the page is, and does it match the page? |
| Content Citability | 6 | 25 | Is the page structured so its claims can be read and attributed? |
| Technical Trust | 3 | 10 | Is the page served securely, quickly and with a 200? |
| Total | 18 | 100 | A = 90–100 · B = 80–89 · C = 65–79 · D = 50–64 · F = <50 |
Category 1 — AI Crawler Access (30 points)
The highest-weighted category, and the only one that is a hard gate: a crawler that is blocked in robots.txt cannot fetch the page, whatever else the page does well. One misplaced Disallow: / removes a site from the live-answer path entirely.
Aura runs this as a single weighted check. Each of the 22 tracked crawlers carries a weight by purpose: live-answer crawlers such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User count most; training crawlers such as GPTBot, ClaudeBot, Google-Extended and CCBot count less. Allow every crawler and you earn all 30 points; each blocked crawler removes its weight.
| # | Check | Kind | Max | How points are awarded |
|---|---|---|---|---|
| 1 | AI crawler permissions (weighted across all tracked bots) | Access blocker | 30 | Proportional to allowed vs. blocked weight. A blocked live-answer crawler is a fail; blocked training-only crawlers are a warning. |
Fix: remove Disallow: / from the AI-crawler entries in robots.txt. Blocking training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) while allowing the live-answer crawlers is a legitimate choice; Aura flags it as a warning, not a fail. Note that GPTBot and ClaudeBot are training crawlers, not the ones that answer live questions.
Category 2 — Machine Discovery (15 points)
Three checks for the conventional files that tell any crawler what a site contains. Two of them are optional conventions and are labelled as such.
| # | Check | Kind | Points | Fix if failing |
|---|---|---|---|---|
| 2 | llms.txt file present | Optional convention | 8 | Optional: publish /llms.txt with the site's purpose and key pages, one line each. The free generator drafts one from your homepage and sitemap. |
| 3 | XML sitemap present | Best practice | 4 | Publish /sitemap.xml and reference it in robots.txt. |
| 4 | RSS/Atom feed present | Optional convention | 3 | Optional: expose a feed for articles and updates. N/A when the site links to no blog or news section. |
llms.txt carries 8 of the 15 points because it is the check sites most often fail and the one most directly written for language-model tools. That weight expresses how much publishing the file changes what such a tool can do with the site; it is not a measured effect on citations. Google says its Search does not use the file, and no major engine has committed to reading it. Publish one because it is cheap and harmless, and edit it before you ship it.
Category 3 — Structured Data (20 points)
JSON-LD markup describes what a page is: an organisation, an article, a set of questions and answers. Five checks cover validity, identity, page type, previews and canonical URL. Since engine 1.2.0, FAQPage markup only earns full points when every question it declares is visible on the page; markup that describes content visitors cannot see is against Google's structured-data guidance and is flagged.
| # | Check | Kind | Max | Notes |
|---|---|---|---|---|
| 5 | JSON-LD structured data (valid) | Best practice | 8 | 8 if present and parse-error-free; 4 if present with errors; 0 if absent. |
| 6 | Organization or WebSite schema | Best practice | 4 | Entity identity: name, URL, logo, official profiles. |
| 7 | Article or FAQPage schema | Best practice | 4 | Only expected on article URLs or pages with FAQ-shaped headings; N/A elsewhere. FAQPage whose questions are not on the page scores 2 with a warning. |
| 8 | Open Graph tags (og:title + og:description) | Best practice | 2 | 2 if both present; 1 if partial; 0 if none. |
| 9 | Canonical URL tag | Best practice | 2 | <link rel="canonical"> prevents duplicate-content dilution. |
Category 4 — Content Citability (25 points)
Six checks on whether a page is structured so that a discrete claim can be found, read and attributed. These are the ordinary conditions of a page that is meant to be read; they are best practices, and none of them guarantees a citation.
| # | Check | Kind | Max | Pass condition / partial credit |
|---|---|---|---|---|
| 10 | Page title length (15–70 characters) | Best practice | 5 | 5 if 15–70 chars; 2 if present but outside the range; 0 if missing. |
| 11 | Meta description length (50–170 characters) | Best practice | 4 | 4 if 50–170 chars; 2 if present but outside the range; 0 if missing. |
| 12 | Single H1 | Best practice | 4 | 4 for exactly one H1; 2 for several; 0 for none. |
| 13 | Heading hierarchy (≥2 H2s) | Best practice | 4 | 4 for two or more H2s; 2 for one; 0 for none. |
| 14 | Quotable blocks (lists and tables) | Best practice | 4 | 4 for two or more; 2 for one; 0 for none. |
| 15 | Content depth (word count) | Best practice | 4 | 4 for 300+ words; 2 for 120–299; 0 below 120. |
Category 5 — Technical Trust (10 points)
Three basic conditions for a page to be fetched and trusted at all.
| # | Check | Kind | Max | Notes |
|---|---|---|---|---|
| 16 | Served over HTTPS | Access blocker | 4 | The URL scheme is checked; the certificate itself is not validated. Non-HTTPS is a fail. |
| 17 | Response time | Access blocker | 3 | 3 under 1.5 s; 2 under 4 s; 1 otherwise. Crawlers time out on slow origins. Single-request timing, not a Core Web Vitals score. |
| 18 | HTTP 200 at the final URL | Access blocker | 3 | The final URL after redirects must return 200. |
How to use this checklist
Start with anything labelled an access blocker: a blocked live-answer crawler costs a large share of 30 points and the fix is a one-line robots.txt change. Then take the best practices in Category 4; a thin page with no heading structure cannot be helped by structured data. Treat the optional conventions last, and only where they describe real content. Every report ranks its fixes by points recovered divided by effort under “Start here”, and every fix says how to validate it: re-run the scan.
Run the free Aura scan to see which of the 18 checks your site passes or fails, with the evidence and the fix for each. The full weighting rationale, the N/A rules and the version history are in the scoring notes that ship with the code.
Limitations
- The scan fetches one page (the URL you enter, following redirects) plus robots.txt, the sitemap and llms.txt. It does not crawl the site.
- Heuristic thresholds (title length, 300 words, two H2s) are editorial choices, not measured optima.
- Response time is one request from one location; it is not a Core Web Vitals or field-performance score.
- Nothing here measures how often any engine mentions or cites the site. That is a sampled measurement with its own uncertainty.
Sources
- llmstxt.org — the llms.txt proposal.
- Google Search Central: AI optimization guide — Google's statements on llms.txt and schema for its AI features.
- Google Search Central: AI features and your website — structured data must match visible content.
- Google Search Central: robots.txt introduction.
Frequently asked questions
How does Aura score AI crawler access?
Aura reads robots.txt for 22 AI crawlers — GPTBot, OAI-SearchBot, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, CCBot and others — each weighted by purpose: crawlers that fetch pages to answer live questions count most, training-only crawlers less. The score is proportional: allow every crawler and you earn all 30 points. Blocking a live-answer crawler is a fail; blocking only training crawlers is a warning, because that is a legitimate choice.
What is llms.txt and does it affect the score?
llms.txt is a proposed convention (llmstxt.org): a plain-text file at your domain root that gives language-model tools a curated summary of the site and its key pages. Aura awards 8 of the 15 Machine Discovery points for it and labels the check an optional convention. Some tools read it; Google states its Search does not use it, and no major engine has committed to it. Publishing one is cheap and harmless, not a ranking lever.
How is the 0-100 score calculated?
Points earned over points available across five categories: AI Crawler Access (30), Machine Discovery (15), Structured Data (20), Content Citability (25) and Technical Trust (10). Each of the 18 checks carries fixed points; checks that do not apply to a page are left out of its denominator. Grades: A = 90-100, B = 80-89, C = 65-79, D = 50-64, F = below 50.
Does a high score mean AI engines will cite my site?
No. The score is a technical checklist of whether AI systems can fetch, read and attribute the site. It does not predict how often any engine mentions or links it. Measuring that takes repeated sampling of real answers, which is what Monitor does, and even then the result is a rate with a confidence interval, not a promise.