Internal Linking Strategy: A Practical Guide
Published · 9 min read
Internal links are the links between pages on your own domain. They do three jobs: they let crawlers discover pages, their anchor text tells crawlers what the target page is about, and they distribute whatever authority the site has toward the pages you link to most. A site that gets those three right has solved most of what people call “site architecture.”
What internal links actually do
Discovery. Google finds most pages by following links from pages it already knows. Google's documentation on crawlable links is explicit that Googlebot follows <a> elements with an href attribute, and that links built any other way may not be followed. A sitemap helps, but Google describes it as a hint about which URLs exist, not a substitute for links.
Context. Anchor text is a short description of the destination written by someone who knows both pages. The same Google page asks for anchor text that gives context about the linked page. If twenty pages link to one URL with the anchor “XML sitemap guide,” Google has twenty consistent statements about what that URL covers before it reads a word of it.
Distribution of authority. PageRank passes through links. Whatever standing your site has from outside links flows inward along your internal links; pages you link to often get more of it, pages you never link to get none. The industry nickname is “link juice.” No one outside Google can tell you how much it is worth on a given query, so treat distribution as a reason to be deliberate, not a lever to tune.
Hub-and-spoke structure
The simplest structure that scales is a hub page for each broad topic, with spoke pages for the subtopics. The hub links to every spoke. Every spoke links back to the hub and, where the content genuinely relates, to sibling spokes. Other names for the same idea are topic clusters and pillar pages.
/technical-seo/ <- hub
/technical-seo/title-tags/ <- spoke, links to hub + siblings
/technical-seo/canonical/ <- spoke
/technical-seo/sitemap/ <- spoke
/technical-seo/core-web-vitals/
/technical-seo/internal-links/URL paths do not have to mirror the structure; the links do. What it buys you: every spoke is two clicks from anything else in the cluster; a new spoke gets a link the day it is published, because adding it to the hub is part of publishing; and the anchor text for each spoke is consistent, because the hub writes it once. On this site the technical SEO checklist plays the hub role for the SEO articles, and the sitemap, canonical, and heading guides are spokes.
For a small site, one hub is enough. For a large site, hubs nest. The test at any size is the same: pick a page at random and count the clicks from the homepage. If the answer is regularly five or more, the structure has a gap.
Orphan pages and how to find them
An orphan page is a page no other page on the site links to. It may be in the sitemap. It may even be indexed, because Google found it once through the sitemap or an outside link. But it gets no anchor text and no distributed authority. Orphans accumulate quietly: old campaign landing pages, posts whose category was deleted, products whose parent went out of stock. Two methods find them, and they work best together.
Crawl versus sitemap diff. Run a crawler from your homepage, following internal links only. Screaming Frog's free tier, Sitebulb, or any crawler you already have will do. Export the list of crawled URLs. Then export every URL in your XML sitemap, or better, every published URL from your CMS. Any URL in the second list but not in the first is an orphan. Most crawlers will do this diff for you if you point them at the sitemap as well.
Search Console. The Links report lists your top internally linked pages. Sort ascending: pages with zero or one internal link are orphans or near-orphans. In the Page indexing report, URLs under “Discovered – currently not indexed” are often pages Google knows only from the sitemap. Search Console lags behind what is live, so treat it as confirmation rather than the primary source.
Aura's free scan checks a single page: robots.txt access for 22 AI crawlers, sitemap and llms.txt presence, structured data, and on-page basics. It does not crawl the whole site, so it cannot find orphans. That is a crawler's job.
Anchor text that describes the target
Good anchor text answers “what will I see if I click this?” in a few words. Three common failure modes:
- Generic anchors. “Click here,” “read more,” “this post.” The link still passes discovery and authority, but it passes no context. Google's crawlable-links documentation specifically lists these as examples of what to avoid.
- Over-optimized exact match. Every link to a page uses the identical keyword-stuffed phrase, in text where no human would write it that way. Google's spam policies cover manipulative link patterns; in our reading, a page whose incoming anchors are all one exact phrase looks manufactured. Vary the wording the way a writer would.
- Ambiguous anchors. The same anchor pointing at two different pages.
<!-- weak: no context -->
For more on canonical tags, <a href="/blog/canonical-tags-guide/">click here</a>.
<!-- weak: same stuffed phrase on every link -->
Read our <a href="/blog/canonical-tags-guide/">canonical tags guide SEO canonical tag best practices</a>.
<!-- good: describes the destination, reads naturally -->
If two URLs serve the same content, set a
<a href="/blog/canonical-tags-guide/">canonical tag pointing at the version you want indexed</a>.Navigation and footer links versus in-body links
Site-wide links, the header navigation and the footer, appear on every page. They are excellent for discovery: anything in the main nav is one click from everywhere. They are weak for context, because the anchor is a short label and the surrounding text is the same everywhere. In our reading, a link repeated on ten thousand pages carries less information per occurrence than one that appears once, in a sentence, on a related page.
In-body links are the opposite. The author decided the reader needed the other page at that moment, the sentence around the link explains why, and the anchor text was written for that spot. They are the links that build topic clusters. A site whose only internal links are in the nav and footer has a flat, uninformative graph, however many links it counts.
The practical split: hubs and the few pages every visitor needs go in the navigation; the footer holds what people actually use there; the real linking work happens inside the content, plus a related-posts module driven by topic, not recency.
A repeatable monthly process
- Export every published URL from the CMS. This is the source of truth, not the sitemap.
- Crawl the site from the homepage, internal links only, and export the crawled URL list.
- Diff the two lists. Every URL in the export but not in the crawl goes on the orphan list.
- For each orphan, decide: link it from its hub and related pages, redirect it, or delete it. A page not worth linking is usually not worth keeping.
- For every page published this month, confirm it is linked from its hub and from the closest related older pages, with descriptive anchors.
- From the crawl, pull internal links that resolve to a redirect or a 4xx status and update them to the final URL.
- Open the Search Console Links report and check that your commercially important pages sit near the top of the internally linked list. If one is near the bottom, it needs hub and nav links, not just one post.
- Record the counts: total pages, orphans, redirected links, broken links. Next month you compare against them instead of guessing.
This fits inside the one-hour SEO audit once the crawler is set up; the first run is the slow one.
Common mistakes and fixes
| Mistake | Why it hurts | Fix |
|---|---|---|
| Internal links pointing at redirected URLs | Extra hop on every crawl; old URL keeps collecting links the new one should get | Update every internal link to the final URL; keep the redirect for outside links |
rel="nofollow" on internal links | Tells Google to discount discovery and context from your own site | Remove it; use noindex on pages you do not want indexed |
| Hundreds of links in the footer | Boilerplate links carry little context; the footer becomes the whole architecture | Trim the footer to what visitors use; move the rest to hub and category pages |
| Pages five or more clicks from the homepage | Crawled less often, little authority reaches them | Add hub pages, category links, and a topic-based related-content module |
| Generic anchor text | No context passed; Google names this as a pattern to avoid | Rewrite anchors to describe the destination page |
| Linking to non-canonical URLs | Signals split between the http, trailing-slash, or parameter variants | Link the canonical URL; see the canonical tags guide |
| Links only rendered by JavaScript after interaction | Google documents that it may not follow links that are not <a href> | Render real anchors in the HTML |
AI systems follow the same graph
AI crawlers reach pages the same way Googlebot does: they fetch a URL, parse the HTML, and follow the <a href> links they find. GPTBot, ClaudeBot, and PerplexityBot have no private map of your site. A page no internal link reaches is as invisible to them as it is to Google, and a hub with consistent anchor text is as clear a statement of what your site covers to a retrieval system as to a search engine. The vendors publish far less about ranking than Google does, so we will not claim to know how they weight links; the discovery mechanism is the same.
Two things complement the link graph without replacing it. The crawlers have to be allowed in, which is a robots.txt question. And an llms.txt file is a curated list of the pages you most want a model to read. It is not a substitute for links: a page listed in llms.txt but linked from nowhere is still an orphan to every crawler that does not read that file.
Frequently asked questions
What is an orphan page?
An orphan page is a page on your site that no other page on your site links to. It may still be in your sitemap and may still be indexed, but crawlers cannot reach it by following links, and it receives no internal anchor text or authority. You find orphans by crawling the site from the homepage and comparing the crawled URL list against your sitemap or CMS export.
How many internal links should a page have?
Google does not publish a target number, and we will not invent one. The useful rule is structural: every important page should be reachable within a few clicks of the homepage, every page should link to its hub and to closely related pages, and each link should be there because a reader might follow it. If a link exists only to hit a count, it is not helping.
Should internal links use nofollow?
No. Google's documentation describes nofollow as a way to qualify links you do not want to vouch for or do not want crawled, which is meant for outbound links. Putting it on your own internal links tells Google to discount the discovery and context those links provide. If you do not want a page indexed, use a noindex tag on that page instead.
Do internal links matter for AI visibility?
Yes, for the same reason they matter for Google. AI crawlers such as GPTBot, ClaudeBot, and PerplexityBot discover pages by following links, and the anchor text and link structure are among the signals a retrieval system has for what a site covers. A page no internal link points to is as hard for an AI crawler to find as it is for Googlebot.