BLOG · SEO

Technical SEO Checklist: 25 Checks to Run Before Anything Else

Published · 9 min read

Technical SEO is the set of conditions that let a search engine fetch a page, understand it, and keep it in the index. If any of those conditions fail, content and links do not matter, because the page is never evaluated. These 25 checks cover the conditions that fail most often, and every one can be verified for free with Google Search Console, PageSpeed Insights, curl, or view-source.

The checks are grouped in the order they should be run. A page that is blocked in robots.txt cannot be helped by a faster server, so crawlability comes first, indexability second, and everything else after.

Summary

CategoryChecksFree tool
Crawlability1–5curl, Search Console robots.txt report, Crawl stats
Indexability6–10view-source, curl -I, Search Console Pages report
URLs & canonicals11–15view-source, curl -IL, URL Inspection
Speed & rendering16–19PageSpeed Insights, Core Web Vitals report, URL Inspection
Mobile & accessibility20–22view-source, curl with a mobile user agent, PageSpeed Insights
Structured data & discovery files23–25Rich Results Test, Schema Markup Validator, Search Console Sitemaps report

Crawlability

1. robots.txt returns HTTP 200. Run curl -sI https://example.com/robots.txt. Google's robots.txt documentation states that a 404 is treated as if no file existed, so everything is crawlable, while a 5xx response is treated far more cautiously. A file that intermittently returns a server error is worse than no file at all.

2. No accidental site-wide disallow. Read the file with curl -s https://example.com/robots.txt and look for Disallow: / under User-agent: * or under Googlebot. This is the single most common way a staging configuration reaches production.

3. The pages that matter are allowed. Open Search Console, use URL Inspection on the home page and one page from each template, and confirm the crawl is allowed. The Settings section also has a robots.txt report that shows the version Google last fetched and any parse errors. If you also care about AI crawlers, the same file governs them; see how to check robots.txt for AI crawlers.

4. CSS and JavaScript are not blocked. Google renders pages before indexing them. If /assets/ or /_next/ is disallowed, Google sees an unstyled or empty page. URL Inspection, then Test live URL, then View tested page, lists blocked resources under More info.

5. No infinite URL spaces. Faceted navigation, calendars, sort parameters, and session IDs can generate millions of crawlable URLs. The Crawl stats report in Search Console (under Settings) shows what Googlebot actually spends requests on. If parameter URLs dominate, block them in robots.txt or remove the links.

Indexability

6. No stray noindex meta tag. View the source of each template and search for robots. A leftover <meta name="robots" content="noindex"> from a launch checklist is common. Google's robots meta tag documentation also notes that a page must be crawlable for the tag to be seen, so do not combine noindex with a robots.txt block.

7. No X-Robots-Tag: noindex header. The same directive can be sent by the server, where view-source will not show it. Check with:

curl -sI https://example.com/page/ | grep -i x-robots-tag

8. The final URL returns 200. Not 200 after a JavaScript redirect, and not a 200 that renders an error message. curl -sIL prints every hop; the last status line is the one that counts.

9. Indexed count roughly matches expected count. The Pages report in Search Console splits URLs into indexed and not indexed, with a reason for each. Compare the indexed number to the number of URLs in your sitemap; a large gap in either direction tells you where to look next.

10. No soft 404s or thin duplicates. The same report lists Soft 404, Duplicate without user-selected canonical, and Crawled – currently not indexed. Each is a list of URLs Google fetched and decided not to keep. Fix the template, not the individual URLs.

URLs & canonicals

11. Every indexable page has a self-referencing canonical. View-source and search for rel="canonical". There should be exactly one, and on a page that is the primary version it should point to itself. Google's canonicalization documentation treats the tag as a strong hint, not a command, so it must agree with the other signals. Failure modes are covered in the canonical tags guide.

12. Canonicals are absolute and use the live host. href="/page/" is technically allowed but fragile. A canonical that still points at staging.example.com or at http:// after a migration tells Google the primary copy is somewhere else.

13. HTTP redirects to HTTPS in one hop, and one host wins. Test all four combinations:

curl -sIL http://example.com/ | grep -iE "^(HTTP|location)"
curl -sIL http://www.example.com/ | grep -iE "^(HTTP|location)"
curl -sIL https://example.com/ | grep -iE "^(HTTP|location)"
curl -sIL https://www.example.com/ | grep -iE "^(HTTP|location)"

Three of the four should be a single 301 to the fourth, which returns 200.

14. Trailing slash and case are consistent. /page and /page/ should not both return 200 with different or identical content. Pick one, redirect the other, and make internal links and the sitemap use the chosen form.

15. No redirect chains or loops. Google's HTTP status code documentation says Googlebot follows up to 10 hops before giving up, but each hop wastes a request and chains are what accumulate over years of migrations. curl -sIL on old URLs from a previous site structure is the quickest way to find them.

Speed & rendering

16. Core Web Vitals pass at the 75th percentile. Google's Core Web Vitals documentation gives the thresholds:

MetricGoodNeeds improvementPoor
Largest Contentful Paint (LCP)≤ 2.5 s2.5–4.0 s> 4.0 s
Interaction to Next Paint (INP)≤ 200 ms200–500 ms> 500 ms
Cumulative Layout Shift (CLS)≤ 0.10.1–0.25> 0.25

PageSpeed Insights shows field data from real Chrome users when enough exists, plus a lab run. The Core Web Vitals report in Search Console groups URLs by status. What to do about each metric is in the Core Web Vitals guide.

17. Server response time is not the bottleneck. The PageSpeed Insights diagnostics section flags slow initial server response. If the HTML itself takes more than a few hundred milliseconds to start arriving, nothing downstream can fix LCP. Caching at the edge or at the application layer is the usual answer.

18. Primary content is in the HTML response. Fetch the page with curl -s and search for a sentence from the main body. If it is absent, the content is injected by JavaScript after load. Google can render JavaScript, but rendering is queued separately and other crawlers may not do it at all.

19. The rendered page matches the raw one. URL Inspection, then View crawled page, shows the HTML Google produced after rendering, along with console errors and blocked resources. If the rendered screenshot is blank or missing sections, that is what Google is indexing.

Mobile & accessibility

20. A viewport meta tag is present. View-source for <meta name="viewport" content="width=device-width, initial-scale=1">. Without it, mobile browsers render the desktop layout scaled down, and layout-based metrics degrade.

21. Mobile and desktop serve the same content. Google's mobile-first indexing documentation states that the mobile version of a page is what Google indexes. Fetch the page with a mobile user agent and compare word count, headings, and structured data to the desktop response:

curl -s -A "Mozilla/5.0 (Linux; Android 10) AppleWebKit/537.36 Chrome/120 Mobile Safari/537.36" https://example.com/page/ | wc -w

22. Basic accessibility passes. PageSpeed Insights includes a Lighthouse accessibility audit that checks tap target size, font size, color contrast, and image alt text. These are not documented ranking factors, but they are a free proxy for whether a page is usable on a phone.

Structured data & discovery files

23. JSON-LD parses and declares who publishes the site. Paste the URL into Google's Rich Results Test or the Schema Markup Validator. A single syntax error invalidates the whole block. The home page should carry Organization and WebSite; which types earn rich results is covered in schema markup for AI search.

24. Content-page schema matches visible text. Article headlines and dates should match the page. FAQPage questions and answers must be on the page word for word. Google's structured data guidelines prohibit marking up content that is not visible to users, and it is an easy thing to drift after an edit.

25. Discovery files exist and are referenced. An XML sitemap should return 200, be listed in robots.txt with a Sitemap: line, and be submitted in Search Console, whose Sitemaps report shows the discovered URL count and any errors. Google's sitemap documentation caps a single file at 50,000 URLs and 50 MB uncompressed; beyond that, use a sitemap index. Only canonical, 200-status URLs belong in it. An RSS or Atom feed and an llms.txt file are optional, but if you publish them, reference them the same way. See the XML sitemap guide and what llms.txt is for the details.

What the Aura free scan does and does not cover

The free scan at the top of this site runs against one URL and covers part of this list: it reads robots.txt for 22 AI crawlers (check 1 in passing, not Googlebot rules), confirms HTTPS and a 200 at the final URL (8, 13), checks that a canonical tag is present (11), validates JSON-LD and looks for Organization, WebSite, Article, and FAQPage (23, 24), and checks whether a sitemap, feed, and llms.txt are present (25). It also samples response time once, which is not a Core Web Vitals measurement.

It does not check for noindex directives, follow redirect chains, inspect blocked resources, measure LCP, INP, or CLS, test a mobile user agent, or crawl any page other than the one you give it. For those, use Search Console and the commands above. If you want a timed walk through the full list, the one-hour SEO audit orders them by payoff.

Frequently asked questions

What is the difference between crawlability and indexability?
Crawlability is whether a search engine is allowed to fetch a URL at all, which is governed mainly by robots.txt and by whether the server responds. Indexability is whether a fetched page is allowed into the index, which is governed by noindex directives, canonical tags, and the HTTP status code. A page can be crawlable but not indexable (a noindex page), and a page can be indexable but never crawled (blocked in robots.txt).

Do I need paid tools to run a technical SEO audit?
Not for these 25 checks. Every one of them can be verified with Google Search Console, PageSpeed Insights, the Rich Results Test, curl, and your browser's view-source. Paid crawlers save time on large sites because they run the same checks across thousands of URLs at once, but they do not see anything Search Console and a terminal cannot show you for a single page.

How often should I run a technical SEO checklist?
Run the full list once, then again after any change to hosting, CMS, templates, or the robots.txt file. Between those events, the Search Console Pages, Sitemaps, and Core Web Vitals reports are enough to catch regressions, so a monthly glance at those three reports covers most sites.

Does the Aura free scan replace a technical SEO audit?
No. The scan looks at one page and covers a subset of this list: robots.txt rules for AI crawlers, HTTPS, the HTTP status at the final URL, the canonical tag, JSON-LD validity and the presence of Organization, WebSite, Article, and FAQPage schema, sitemap, llms.txt, and feed presence, and a single response-time sample. It does not measure Core Web Vitals, check noindex directives, follow redirect chains, test mobile rendering, or crawl the rest of the site.

Run the free 18-check scan on one page

Technical SEO Checklist: 25 Checks (2026) | Aura AI Visibility