Internal Links Are a Directed Retrieval Graph

A graph-first method for crawlable links, reachability, depth, contextual edges, orphan diagnosis, and repair plans that preserve user intent.

Direct answer: internal linking graph

An internal linking graph models pages as nodes and crawlable links as directed edges so teams can measure depth, orphan risk, inlink concentration, anchor context, cluster connectivity, and the actual retrieval paths supporting each important page.

Original research artifact

A directed-edge schema, breadth-first depth calculation, orphan test, anchor audit, and cluster-connectivity checklist.

What this page adds

Connect graph measures to user journeys and crawl evidence instead of presenting PageRank-like scores without inspectable edges.

Related research

Memo Details

Category: SITE ARCHITECTURE. Author: SULAYMAN BOWLES. Published: 2026.07.19. Read time: 13 MIN. Source count: 5.

Evidence Boundary

Graph metrics describe the captured internal architecture and help prioritize review. They do not measure proprietary ranking weights, guarantee crawling or indexing, or prove that adding a link will improve search performance.

Article Metrics

Graph

DIRECTED

Edge context

RETAINED

Orphan states

04

Primary artifact

LINK EDGE TABLE

Research Note

An internal link is a directed edge from one document to another, observed in a particular source or rendered state with anchor text, placement, and crawlability conditions. Reducing that edge to a destination count discards the information needed to explain architecture. A global navigation link, an editorial citation, a related-product card, a pagination control, and a hidden script route can all point to the same URL while serving different retrieval functions.

The graph-first approach preserves node identity and edge context, then asks reachable questions: which canonical pages can be reached from approved entry points, how many hops are required, which components are disconnected, where do redirects consume edges, and which important pages depend on one fragile source? Metrics narrow the review set. They do not replace the product judgment that determines which paths should exist.

Store a link edge with enough context to review it

The minimum edge has a source URL, observed href, resolved destination, anchor or accessible name, DOM or source location, link type, discovery state, and artifact reference. Preserve whether the element was an anchor with an href because Google documents that form as the generally crawlable link pattern. A click handler on a div may support a user interaction while remaining a different technical object.

Normalize the source and destination against the same route-identity contract used by the crawl frontier, but retain the observed href. Relative paths, fragments, alternate hosts, tracking parameters, and redirects can then be diagnosed without losing evidence. Classify the destination after resolution as canonical, alias, external, asset, non-HTTP, malformed, or unknown.

Placement should be derived cautiously. Header, primary navigation, breadcrumb, main content, related content, pagination, filter, and footer are useful classes when the DOM structure supports them. Do not infer importance from CSS coordinates alone. The important point is that a repeated template edge and a unique editorial edge remain distinguishable.

Internal-link edge contract
FieldExampleReview question
source_key/guides/crawlingWhich canonical page emits the edge?
observed_href../tools/crawler?ref=guideWhat did the document actually contain?
destination_key/tools/crawlerWhich canonical identity receives it?
labelcrawler architecture toolCan a reader predict the destination?
placementmain/editorialIs the edge contextual or repeated template chrome?
artifact_refrun/url/source#node-184Can another reviewer reproduce the observation?

Measure reachability from real entry points

A page is structurally reachable when a directed path exists from an approved entry set through crawlable internal edges. The homepage is one entry, not the only one. Category hubs, section roots, locale roots, authenticated application shells, and feed or campaign landings can represent different user journeys. Define the entry set before calculating depth.

Run breadth-first search on canonical destinations after resolving known redirects. Report shortest depth and at least one path, but also count independent parent sources. A priority page reachable in two hops through one soon-to-be-retired article is more fragile than a page with several relevant parents. Path diversity is not a ranking metric; it is an architecture resilience measure.

Keep source and rendered graphs separate when client execution adds navigation. If the source graph cannot reach an article but the rendered graph can, the finding is dependency on the tested client path, not absolute orphaning. If neither graph reaches it but a sitemap lists it, label it sitemap-only. Each state implies a different repair.

Reachability states for a public canonical page
StateObserved evidenceInterpretation
source-reachablePath through source anchor href edgesServer-delivered discovery path exists
render-onlyNo source path; path appears after tested executionDiscovery depends on client rendering scenario
sitemap-onlyNo internal path; canonical URL is nominated in sitemapDeclared but not architecturally integrated
external-onlyKnown inbound reference but no internal pathThird-party discovery exists; internal context absent
unobservedNo captured path or nominationCoverage or publication gap; verify inventory first

Diagnose components before labeling orphan pages

A directed graph can contain weakly connected groups that are isolated from the primary site and strongly connected groups that circulate internally without a path from an entry. These components often reveal retired subdirectories, alternate hosts, faceted islands, staging remnants, or microsites imported without navigation. Inspect representative URLs and edge classes before assigning one site-wide remedy.

An orphan label needs an inventory source. A page discovered only through a database export, analytics, log file, external backlink, or sitemap is not equivalent to a page discovered during the internal crawl. Record the source that proves the node exists. Without that, the absence of inlinks may simply mean the crawler never knew the URL.

Not every zero-inlink URL needs a new link. Redirect targets, campaign pages, utility endpoints, legal notices, feed documents, and intentionally private states may have special entry paths or should leave the indexable inventory. The decision should be keep and link, consolidate and redirect, retain with bounded reachability, noindex, or remove—not add links everywhere.

  • Compare crawl nodes with sitemap, analytics, logs, CMS exports, and approved route registries.
  • Separate zero observed inlinks from zero eligible crawlable inlinks.
  • Inspect weak and strong components for shared template or host causes.
  • Assign an intended entry path before proposing a link.
  • Remove aliases and noncanonical URLs from the candidate set before counting orphans.

Use graph metrics to find questions, not to manufacture importance

In-degree identifies destinations receiving many captured edges; out-degree identifies sources that distribute many. Depth exposes long retrieval paths. Betweenness can identify bridge pages whose removal disconnects routes. Strong components expose circulation. Edge concentration shows whether most paths depend on one template. These measures are useful because they reveal structures worth reviewing.

Raw PageRank or similar centrality scores can be calculated on an internal graph, but the score is a model of the captured link network, not a disclosed search-engine value. Results are sensitive to which nodes and edges were included, how redirects were resolved, whether repeated links were collapsed, and how dangling nodes were handled. Store those choices next to the output.

Anchor text and surrounding context answer a different question: whether the relationship is legible. Repeated exact-match anchors are not automatically better, and generic labels are not automatically wrong. Review whether the phrase helps a reader understand the destination, fits the sentence or component, and distinguishes nearby links.

Graph metrics as diagnostic prompts
MetricUseful questionCommon misuse
shortest depthHow many eligible hops from the intended entry?Treating every deep page as low quality
in-degreeWhich destinations attract or lack internal references?Equating count with proprietary authority
parent diversityHow many independent sources support the route?Counting repeated template links as independent
betweennessWhich pages bridge otherwise separate sections?Changing bridges without journey review
component membershipWhich routes circulate outside the main architecture?Calling every isolated component an error

Repair journeys, then rerun the same graph contract

A link plan should name the source page, destination canonical, placement, intended reader question, anchor or component label, owner, and expected graph change. This prevents the recommendation “add more internal links” from becoming a site-wide template block. The strongest repairs usually connect an existing explanatory page to the next useful artifact or decision.

After implementation, crawl the same scope and rebuild both source and rendered edge tables. Assert that the new edge is a crawlable anchor, reaches the canonical destination without an avoidable redirect, appears in the intended component, and does not create broken or duplicate links. Recalculate reachability and depth only after the edge-level checks pass.

Monitor for regressions when templates or route registries change. A page can return to sitemap-only status when a collection stops rendering, a navigation key changes, or a canonical path is renamed without updating contextual links. The durable artifact is therefore not a one-time score; it is a versioned graph and a small set of architecture invariants.

Find sitemap URLs with no eligible internal inlinks — This query assumes redirects and canonical aliases were resolved before edge loading.
SELECT u.canonical_url
FROM url_inventory AS u
LEFT JOIN link_edges AS e
  ON e.destination_key = u.url_key
 AND e.is_crawlable = 1
WHERE u.in_sitemap = 1
  AND u.is_indexable = 1
GROUP BY u.url_key, u.canonical_url
HAVING COUNT(e.source_key) = 0
ORDER BY u.canonical_url;

Thesis

Internal-link quality is the ability of people and crawlers to reach the right canonical pages through meaningful, crawlable paths—not the number of links a template can emit.

Source Ledger

Internal Links