Skip to content
siterank.info

Content strategy

Internal Linking for SEO, Retrieval and AI Discovery

What Google actually says internal links do, how anchor text carries meaning, how to find orphan pages, and which retrieval claims are evidence-backed and which are not.

By ankitkumarvig@gmail.com Published September 9, 2026 Last reviewed September 9, 2026 Sources verified September 9, 2026 10 min read

Internal linking is the cheapest structural work on a website and the most commonly neglected. It costs nothing but editorial attention, it is entirely under the site owner's control, and unlike almost everything else discussed under the heading of AI discoverability, Google documents what it does. This article separates the part that is stated in Google's own documentation from the part that is inference about retrieval systems, and shows how to find the pages on your own site that the structure has forgotten.

Google's link best practices page makes two claims about purpose, and they are worth quoting precisely because so much internal-linking advice drifts past them: "Google uses links as a signal when determining the relevancy of pages and to find new pages to crawl."

Two jobs, then. Discovery — a crawler follows links to reach URLs it has not seen. The SEO starter guide is blunter: "the vast majority of the new pages Google finds every day are through links." Relevance — the link, and the words used to make it, tell Google something about the destination.

Both are OFFICIAL PROVIDER GUIDANCE. Neither is a claim about how much a link is worth, and Google publishes no such number. What follows from the two documented jobs is a small set of practical rules that are not negotiable.

Google states the format requirement directly: "Google can only crawl your link if it's an <a> HTML element with an href attribute." A <span> with a click handler, a <div> styled to look like a link, a button that pushes a route, or an href populated only after a JavaScript event — none of these are reliably a link as far as crawling is concerned.

This matters more on WordPress than people expect. Page builders, tabbed layouts, card grids and "load more" patterns routinely produce navigational elements that are not anchors. If the only path to a page is through one of those, the page is effectively undiscovered by links, whatever it looks like in a browser.

OFFICIAL PROVIDER GUIDANCE. This is a correctness requirement, not an optimisation.

Anchor text

Google's guidance on link text is short and specific. Anchor text "tells people and Google something about the page you're linking to." It should be descriptive rather than generic — "click here" and "read more" are named as things to avoid — and it should be concise. The page also notes that "the words before and after links matter, so pay attention to the sentence as a whole", which means the surrounding sentence is part of the context, not just the linked words. And it warns against keyword stuffing in anchor text.

The practical version:

Anchor Verdict Why
read more Poor Describes nothing about the destination
entry requirements Good Names the destination's subject in the reader's words
undergraduate entry requirements for International Baccalaureate applicants Acceptable but long Concise is stated guidance; a whole sentence as anchor obscures the surrounding text
payroll software payroll compliance software best payroll software Harmful Keyword stuffing, explicitly warned against
The destination's exact page title, on every link to it Suspicious Identical anchors everywhere reads as templating, not editorial judgement

Vary the anchor naturally. Different pages have different reasons to link to the same destination, and the anchor should reflect the reason.

The single most actionable line in Google's link documentation is this: "Every page you care about should have a link from at least one other page on your site."

A page with no inbound internal link is an orphan. It may still be in the XML sitemap, and Google may still crawl and index it, so an orphan is not automatically invisible. What it lacks is the relevance context that a link provides, and the browsing path that a reader would use. On most WordPress sites orphans are not deliberate; they are the residue of a redesign, an imported archive, a landing page built for a campaign that ended, or a post that never made it into a category.

Two definitions are used in practice, and mixing them produces bad reports:

  • Strict orphan — no inbound link from any source, including menus, breadcrumbs, archives and sitemaps.
  • Content orphan — no inbound link from within the body content of another page, even though a menu or archive template may link to it.

SiteRank AI reports the second, and says so in the finding text. The link graph is built from links in the rendered main content, after navigation, header, footer and sidebar elements are stripped. A page reachable only from the primary menu is therefore reported. That is a deliberate choice: a template link appears on every page and carries no editorial signal about any particular one, so counting it would hide exactly the pages that need attention. It is also a choice a reader of the report needs to know about, which is why the rule states it rather than leaving the number unexplained.

Weak internal linking is different from orphaning

A page with one inbound link on a site whose median page has twelve is not an orphan, but it is under-linked relative to its own site. Fixed thresholds are wrong here — link density varies by an order of magnitude between a 40-page brochure site and a 4,000-post publication — so the comparison has to be internal.

SiteRank's weak-linking rule derives its threshold from the corpus: it considers only indexable pages of at least 300 words that already have at least one inbound link, computes the median inbound count across the site, and flags pages at or below a quarter of that median (with a floor of one). Below ten indexable pages it reports nothing at all, because a median over a handful of documents is noise rather than a benchmark.

The finding shows the page's inbound count and the site median side by side. A number a reader cannot decompose is not a finding; it is an assertion.

Adding links one at a time does not fix a structure. The structure questions are:

Depth. How many clicks from the home page to the page, following content links? Deep pages are crawled less often and are harder for readers to reach. There is no published threshold; the useful test is whether a reader who lands on the home page with a specific question can get there in a few obvious steps.

Hubs. Does each subject have a page that gathers its parts? This is the pillar page idea, covered in detail in pillar pages and topic clusters. A hub concentrates inbound links from its cluster and gives the crawler a single entry point to the subject.

Directories. Google's starter guide notes that "using directories (or folders) to group similar topics can help Google learn how often the URLs in individual directories change." On WordPress this is a permalink and taxonomy decision, and it is far cheaper to make before publishing 500 posts than after.

Body links versus template links. A link in a sidebar widget appears identically on every page. A link written into a sentence exists because the writer judged it relevant there. Both are crawlable; only the second carries editorial meaning, and only the second gets natural, varied anchor text.

Reciprocity within a cluster. Supporting pages link up to the hub; the hub links down to each supporting page at the point where it summarises that facet. Lateral links between supporting pages where two facets genuinely relate. Not everything to everything.

What internal linking does for AI retrieval

Here the evidence thins, and the honest thing is to say where it stops.

The documented part: Google states that to be shown as a supporting link in AI Overviews or AI Mode a page must be indexed and eligible to be shown in Google Search with a snippet, and that there are no additional requirements or special optimisations. Since links are how Google finds pages, internal linking is upstream of AI eligibility for exactly the same reason it is upstream of ordinary indexing. The same AI features page also lists maintaining high-quality internal links among the ordinary practices that apply. OFFICIAL PROVIDER GUIDANCE — for the fact that links matter to indexing, not for any AI-specific effect.

The inferred part: retrieval systems generally operate on passages rather than whole documents, so the passage retrieved from your page is unlikely to carry the surrounding page's navigation with it. Two consequences are plausible and unconfirmed:

  • A page that must be discovered before it can be embedded or indexed is subject to the same discovery mechanics as any other page. EMERGING PRACTICE as a retrieval claim; OFFICIAL PROVIDER GUIDANCE as an indexing claim.
  • Descriptive anchor text on the linking page describes the destination in a way that is independent of the destination's own wording, which is a second phrasing of the same subject. Whether any AI system uses that is undocumented. HYPOTHESIS.

And the part that should not be claimed at all: that internal links transfer authority to a page in a way that makes an AI system more likely to cite it. No provider documents citation selection. Anyone who tells you a link graph feeds a "trust score" inside a language model is describing something they cannot have observed. Crawled is not cited, and the gap between those two states is where most internal-linking-for-AI advice quietly lives.

Evidence summary

Practice Tier Basis
Use <a href> for anything you want crawled OFFICIAL PROVIDER GUIDANCE Google: "Google can only crawl your link if it's an <a> HTML element with an href attribute"
Give every page you care about at least one inbound internal link OFFICIAL PROVIDER GUIDANCE Google link best practices, stated as a rule
Write descriptive, concise anchor text; avoid "click here" and keyword stuffing OFFICIAL PROVIDER GUIDANCE Google link best practices
Pay attention to the sentence around the link OFFICIAL PROVIDER GUIDANCE Google: "the words before and after links matter"
Group similar topics into directories OFFICIAL PROVIDER GUIDANCE Google SEO starter guide
Maintain high-quality internal links as part of AI-feature readiness OFFICIAL PROVIDER GUIDANCE Google AI features documentation, listed among ordinary practices
Hub-and-spoke linking improves performance for the cluster's subject EMERGING PRACTICE Plausible mechanism, widely practised, no provider confirmation of a cluster-level effect
Internal links influence which passages an AI answer engine retrieves HYPOTHESIS No provider documents retrieval selection
Internal links influence which sources an AI answer engine cites HYPOTHESIS No provider documents citation selection

Common failures, and what to do instead

Failure What it looks like Fix
Automated related-posts widgets counted as internal linking Every post links to three algorithmically chosen posts, identically templated Keep the widget if readers use it; do not count it as editorial linking, and add body links
A "link to five other posts" rule Forced links to weakly related pages, generic anchors Link where the sentence genuinely wants a reference
Bulk anchor-text automation The same phrase auto-linked site-wide to the same URL This is the pattern Google's keyword-stuffing warning targets
Linking only from new content Old high-traffic pages never updated to point at newer, better pages Work backwards: find your best-linked pages and add links from them
Redesign amnesia A new theme drops the old contextual links; orphan count triples overnight Crawl after every structural change
Links to non-indexable pages Body links pointing at noindex or canonicalised-away URLs Fix the target or drop the link; a link to a page that cannot be indexed does no discovery work

What not to assume

  • That an orphan page is invisible. It can be crawled from a sitemap. What it lacks is context and a browsing path.
  • That more links is better. Google's own warning about keyword-stuffed anchor text and the general principle that links express editorial judgement both argue against volume. A page linked from everywhere with the same anchor tells a crawler less, not more.
  • That nofollow shapes internal crawling usefully. Google's guidance frames nofollow around untrusted, paid or user-generated links. Using it to sculpt internal crawl paths is not what it is documented for.
  • That a link graph built from a plugin's data equals reality. A stored graph goes stale, misses JavaScript-injected links, and depends on how boilerplate was stripped. Confirm orphan findings by crawling.
  • That fixing internal links will move an AI visibility number. It might; you cannot claim it did without a controlled measurement, and even then the honest word is association, not cause. The methodology page sets out what a defensible before-and-after comparison requires.

Key takeaways

  • Google documents two jobs for links: finding new pages, and signalling relevance. Everything defensible follows from those two.
  • The link must be an <a> element with an href. Anything else is a design element, not a link.
  • Anchor text should be descriptive, concise, varied, and never stuffed. The surrounding sentence counts.
  • Every page you care about needs at least one inbound internal link. Content orphans are common and invisible without a crawl.
  • Compare link counts against your own site's median, not a fixed target.
  • Internal linking is upstream of AI features because it is upstream of indexing. Any claim beyond that about retrieval or citation is a hypothesis.

Official sources & further reading