Industry Report #002

Finance's Invisible Web: The Content Nothing Links To

On ten of twelve Fortune 500 finance sites where it can be measured, an average of 35% of published content has no inbound link anywhere, present only in the XML sitemap, reaching 79% on the worst site. Only two payment networks block the very inventory that would let anyone check.

Industry · Financial Services Dataset · Q2 2026 Sites analyzed · 12 Published · June 23, 2026
TL;DR The fast read. Each line is measured below.
  • 35% of published content is link-invisible on average, up to 79% on one site: live and in the sitemap, but nothing links to it, so a link-following AI agent never finds it.
  • 54% is all that is reachable the way an unaided agent reads a site. The rest is navigation-only or invisible.
  • 2 of 12 sites cannot be measured at all: payment networks that block their own sitemap from crawlers.
  • 5.6% of schema markup is answer-shaped (FAQ / Q&A). The other ~81% is identity and navigation scaffolding: who they are, not what they answer.
  • 45% of a site's structural load runs through its top 1% of pages on average, 92% at the extreme. These webs hang on a few hubs.
  • 27% / 14% orphans versus dead-ends, two independent failures on different sites, so no single fix closes both.

The Headline

Across twelve Fortune 500 financial-services sites crawled in Q2 2026, totaling 10,731 pages and 15.1 million tokens of public content, the surface looks healthy. Underneath runs a fault line that conventional audits miss: a large share of published content that nothing on the site links to, present only in the XML sitemap, invisible to any AI agent that discovers content by following links.

On the ten sites where we can measure this directly, the share of link-invisible content does not sit quietly near zero. It ranges from 2% to 79%, averaging 35%. This is not the work of one or two freak outliers dragging an otherwise clean cohort. It is a broad structural habit, present to some degree on every measurable site and severe on four of them: a card-and-payments brand at 57%, an investment bank at 43%, a brokerage at 74%, and a second investment bank at 79%, where years of research and press archive exist as live URLs that the live site no longer points at.

The other two sites, both payment networks, tell a different version of the problem. On them, invisibility could not be measured at all, because the crawl could not read a complete inventory of their pages. That is its own finding, and we treat it as one rather than papering over it with a zero.

Why Ten Sites, Not Twelve

Link-invisibility is a question about a site’s full inventory: of every page it publishes, how many does nothing link to. Answering it needs that inventory, which is the XML sitemap. Read the sitemap, and you can compare the published set against the link graph to find the unreachable pages. Without it the question is unanswerable, because an agent following links and a crawler reading the sitemap converge on the same visible pages, so anything neither reaches is never seen and cannot be counted.

Two of the twelve, both payment networks, returned no sitemap at all: they block automated sitemap and robots.txt requests behind the same anti-bot layer that guards their login flows. We record those as not-measurable, not zero, because counting them as zero would understate the problem precisely where it is worst. It is a small but telling minority: the institutions that most aggressively wall off their inventory are the payment rails everything else runs on, and to an AI agent their content is simply unlistable.

The rest of this report keeps the two questions separate. Reachability and invisibility are reported only for the ten measurable sites. Topology, fragility, content, and the scorecard use all twelve, because those measures do not depend on the sitemap.

The Three-Tier Reachability Model

Every page on a site falls into one of three reachability tiers, judged against the full link graph (navigation included) for reach and the content link graph for editorial authority:

  1. Reachable via editorial links. An in-content link path leads to the page. These are the pages a link-following AI agent can find on its own.
  2. Navigation-only. The page is reachable, but only through chrome: header, footer, or menu links that repeat site-wide. An agent that filters boilerplate navigation, as most do to find the actual content, loses these.
  3. Link-invisible. The page has zero inbound links of any kind in the full graph, editorial or navigational. It exists as a URL and is listed in the sitemap, but nothing on the site advertises it to a crawler that follows links.

Averaged across the ten measurable finance sites:

Reachable via editorial links: 54.2% Reachable via editorial links 54.2% Navigation-only: 10.6% 10.6% Link-invisible: 35.2% Link-invisible 35.2%
  • Navigation-only (10.6%)
Mean across the 10 measurable finance F500 sites, Q2 2026.
View data
Segment Value
Reachable via editorial links 54.2%
Navigation-only 10.6%
Link-invisible 35.2%

Just over half of published content, on average, is reachable the way an unaided agent reads a site, and the rest is not. That is a starkly different picture from the surface health these sites project, and it is not carried by a single anomaly. Here is how the link-invisible tier distributes across the ten:

0% 20% 40% 60% 80% sample mean 35.2% SF4: 17.2% SF4 SF5: 16.1% SF5 SF9: 16% SF9 SF7: 79.3% SF7 SF8: 2.2% SF8 SF2: 13.4% SF2 SF10: 34.2% SF10 SF11: 42.7% SF11 SF3: 56.7% SF3 SF6: 74.5% SF6 Link-invisible page rate (% of site's pages)
Each dot is one of the 10 measurable sites. Mean 35.2%.
View data
Site Link-invisible page rate (% of site's pages)
SF8 2.2%
SF2 13.4%
SF9 16%
SF5 16.1%
SF4 17.2%
SF10 34.2%
SF11 42.7%
SF3 56.7%
SF6 74.5%
SF7 79.3%
Sample mean 35.2%

Every measurable site has a non-trivial invisible tier, and four of the ten bury between 43% and 79% of their content. These are not degraded crawls. The invisible pages carry real body content and were discovered through the sitemap. The pattern repeats: a large editorial archive, an insights-and-research library on one site, a press and worldwide archive on another, that the live navigation has stopped pointing at. The content was published, then orphaned as the site moved on, and now survives only in the sitemap. For an AI agent, that archive is dark.

The navigation-only tier tells a quieter version of the same story. One site routes 70% of its pages through chrome alone: reachable, but only if the agent trusts the repeated header and footer links it would normally discard. The safe assumption for an institution is that the agent does not.

Machine-Readability: Reaching a Page Is Not Reading It

Reaching a page is one thing; parsing it is another. Two crawl signals decide whether an agent can read what it reaches, and both split the cohort in half.

The first is structured data: the JSON-LD and schema.org markup that states a page’s meaning in machine terms instead of leaving an agent to infer it from prose. Across the twelve sites, 56% of pages carry structured data on average, but the distribution is bimodal, running from under 1% to 100%.

0% 20% 40% 60% 80% 100% sample mean 55.9% SF6: 99.6% SF6 SF8: 3.6% SF8 SF5: 94.8% SF5 SF2: 0.6% SF2 SF3: 47.6% SF3 SF12: 93.7% SF12 SF10: 0.4% SF10 SF1: 18.5% SF1 SF9: 45.7% SF9 SF11: 75.6% SF11 SF4: 91.2% SF4 SF7: 100% SF7 Pages carrying JSON-LD / schema.org structured data (%)
Each dot is one site. Mean 55.9% (n=12). The cohort splits into near-full and near-zero.
View data
Site Pages carrying JSON-LD / schema.org structured data (%)
SF10 0.4%
SF2 0.6%
SF8 3.6%
SF1 18.5%
SF9 45.7%
SF3 47.6%
SF11 75.6%
SF4 91.2%
SF12 93.7%
SF5 94.8%
SF6 99.6%
SF7 100%
Sample mean 55.9%

But coverage hides composition. Sorting every schema.org annotation across the cohort by what it actually describes tells a sharper story than the headline percentage:

Navigation
41.3%
WebPage, BreadcrumbList, ListItem
Identity
40.1%
Organization, ContactPoint, Person
Article / media
11.3%
Article, NewsArticle, VideoObject
Answer (FAQ / Q&A)
5.6%
FAQPage, Question, Answer
Product / other
1.7%
CreditCard, LoanOrCredit

Share of all schema.org annotations across the 12 sites, by type.

22,644 annotations cohort-wide. Over four fifths is navigation and identity scaffolding.

More than four fifths of all structured data is navigation and identity scaffolding: who the company is, where to click. Only 5.6% is answer-shaped, the FAQ and question-and-answer markup an AI assistant can lift straight into a reply. One site marks up 100% of its pages and carries none of it. The institutions are describing who they are to a machine, not what they can answer.

The second signal is render dependency: pages whose content only appears after a browser runs their JavaScript. An agent that fetches raw HTML, which most do by default because rendering every page is expensive, sees an empty shell. On average 20% of pages require rendering, but one payments site requires it for 94%.

0% 20% 40% 60% 80% 100% sample mean 20.4% SF12: 5.5% SF12 SF10: 4.9% SF10 SF11: 3.7% SF11 SF3: 2.7% SF3 SF6: 1.6% SF6 SF8: 0.1% SF8 SF2: 20.3% SF2 SF7: 0% SF7 SF9: 12% SF9 SF5: 44% SF5 SF4: 55.3% SF4 SF1: 94.1% SF1 Pages whose content requires browser rendering (%)
Each dot is one site. Mean 20.4% (n=12). Higher means more content is hidden behind JavaScript.
View data
Site Pages whose content requires browser rendering (%)
SF7 0%
SF8 0.1%
SF6 1.6%
SF3 2.7%
SF11 3.7%
SF10 4.9%
SF12 5.5%
SF9 12%
SF2 20.3%
SF5 44%
SF4 55.3%
SF1 94.1%
Sample mean 20.4%

This is the most overlooked form of invisibility because the site’s own team never sees it: the page looks complete in a browser. To a raw-HTML agent it reads as blank, perfectly linked and still empty. On the site at 94%, almost the whole property is legible only to a crawler willing to render every page, and most will not.

Put together, a finance page is fully visible to an AI agent only when all three hold at once: something links to it, its content survives without JavaScript, and it carries structured data the agent can parse. Few pages in this cohort clear all three bars.

Fragility: The Whole Site Hangs on a Few Pages

The second pattern, concentration, holds across all twelve sites. Betweenness centrality counts how often a page sits on the shortest path between two others: a high-betweenness page is a bridge that traffic, link equity, and an agent’s traversal all funnel through. When that concentrates in a tiny fraction of pages, the site is fragile, structurally dependent on those few and quick to fragment if they are weak or off-topic.

0% 20% 40% 60% 80% sample mean 44.6% SF3: 34.4% SF3 SF1: 28.3% SF1 SF9: 27.2% SF9 SF4: 26.6% SF4 SF12: 38.7% SF12 SF5: 52.4% SF5 SF6: 21% SF6 SF10: 33.4% SF10 SF11: 45.1% SF11 SF2: 55.8% SF2 SF8: 79.8% SF8 SF7: 92.3% SF7 Share of total betweenness carried by the top 1% of pages (%)
Each dot is one site. Mean 44.6% (n=12). Higher means more fragile.
View data
Site Share of total betweenness carried by the top 1% of pages (%)
SF6 21%
SF4 26.6%
SF9 27.2%
SF1 28.3%
SF10 33.4%
SF3 34.4%
SF12 38.7%
SF11 45.1%
SF5 52.4%
SF2 55.8%
SF8 79.8%
SF7 92.3%
Sample mean 44.6%

On the average site, the top 1% of pages carry 45% of all betweenness: the structure that holds the site together routes through a handful of hubs. The most concentrated site pushes 92% through 1% of its pages. On these sites, betweenness pools in locale gateways, corporate-about pages, and contact pages. The product catalog hangs off those hubs rather than carrying the structure itself.

For a link-following agent, this compounds the reachability problem. The agent’s picture of “what this company does” is mediated almost entirely by a few pages. If those pages are thin, jurisdictional, or legal, the agent inherits that framing, and whatever sits in the invisible tail never enters the picture at all.

Health Metrics: The Failure Modes Don’t Align

The conventional intuition is that orphans, pages with no inbound editorial links, are the dominant structural problem. The finance data tells a more interesting story, and it is the same story the healthcare cohort told: the sites with the worst orphan rates are not the same sites with the worst dead-end rates.

0% 20% 40% 60% 80% sample mean 27.1% SF2: 5% SF2 SF11: 37.4% SF11 SF8: 2% SF8 SF5: 37.4% SF5 SF12: 1.2% SF12 SF10: 35.7% SF10 SF1: 0.2% SF1 SF9: 14.2% SF9 SF4: 21.9% SF4 SF6: 34.6% SF6 SF3: 56% SF3 SF7: 79.3% SF7 Orphan rate (% of site's pages)
Pages with zero inbound editorial links. Mean 27.1% (n=12).
View data
Site Orphan rate (% of site's pages)
SF1 0.2%
SF12 1.2%
SF8 2%
SF2 5%
SF9 14.2%
SF4 21.9%
SF6 34.6%
SF10 35.7%
SF5 37.4%
SF11 37.4%
SF3 56%
SF7 79.3%
Sample mean 27.1%
0% 20% 40% sample mean 13.7% SF9: 4% SF9 SF3: 1.3% SF3 SF6: 9.3% SF6 SF1: 22.1% SF1 SF8: 1% SF8 SF4: 6.9% SF4 SF10: 20.2% SF10 SF7: 0% SF7 SF12: 5.1% SF12 SF11: 17.5% SF11 SF2: 25.1% SF2 SF5: 51.5% SF5 Dead-end rate (% of site's pages)
Pages that receive inbound links but link out to nothing. Mean 13.7% (n=12).
View data
Site Dead-end rate (% of site's pages)
SF7 0%
SF8 1%
SF3 1.3%
SF9 4%
SF12 5.1%
SF4 6.9%
SF6 9.3%
SF11 17.5%
SF10 20.2%
SF1 22.1%
SF2 25.1%
SF5 51.5%
Sample mean 13.7%

Compare the two strips. The site with the highest orphan rate (SF7 at 79%) has the lowest dead-end rate in the cohort (0%): it strands its pages at the entrance, but every page it does expose links onward cleanly. The site with the highest dead-end rate (SF5 at 52%) sits mid-pack on orphans. The two means are different, 27% for orphans against 14% for dead-ends, and they fall on different sites: the rank correlation between them is weak. These are independent failure modes, not two faces of the same problem. A site can be excellent at receiving its visitors and terrible at passing them onward, or the reverse.

The orphan strip also shows finance’s signature spread. Half the cohort sits below 22%, then the rate climbs steeply to 56% and 79% on the two sites that also carry the heaviest link-invisible tail. Orphaning at the editorial level and invisibility at the full-graph level are the surface and the depth of one underlying habit: publishing content the navigation has stopped pointing at.

The Industry Scorecard

The five-lens analysis assigns each site a green, amber, or red score on each lens. Across all twelve sites:

Site Skeleton Size, density, and average path length. How big and how connected the site is at the body level. Circulation PageRank distribution and structural bottlenecks. How importance flows between pages, and which hubs hold it all together. Organs Community detection. Whether the site's topical clusters cleanly separate, or whether one mega-cluster dominates everything. Health Islands, orphans, and dead-ends. Where content is structurally dying: unreachable, unlinked, or terminating. Nervous Sys. Click depth, bridges, and cross-community linking. Whether the site is a well-designed building or a pile of disconnected rooms.
SF1
SF2
SF3
SF4
SF5
SF6
SF7
SF8
SF9
SF10
SF11
SF12
  • Green: healthy
  • Amber: moderate concern
  • Red: critical
Five-lens scorecard for each of the 12 finance sites.

The pattern echoes healthcare: skeletons are mostly fine; the organs and how signal travels between them are the soft spots. Ten of twelve sites earn green skeletons and the other two earn amber, so link density at the body level is healthy. But Health is the worst-scoring lens by a wide margin, eight of twelve sites score red, and Organs (topical neighborhoods) is close behind with five reds. The picture is consistent across two industries now: F500 sites are well-built at the surface and broken at the level of how editorial signal travels between topical communities. The four sites carrying the heaviest invisible tail all score red on Health, the lens that captures orphaning and dead-ending. Structural failures compound: the site that hides content from agents also struggles to circulate the content it does expose.

The Density Spectrum

The finance cohort does not have a single shape. Link density spans more than two orders of magnitude, from 178 content edges on the sparsest site to 63,564 on the densest, on a broadly similar number of pages.

PoleEdgesBehavior
Dense mesh22,000–64,000Hyperlinked, shallow, low orphan risk
Sparse tree178–8,000Hierarchical, few cross-links, deep paths, high dead-end risk

Two sites behave like switchboards: nearly every page links to dozens of others, and their orphan rates are among the lowest in the cohort. The rest sit in tree-like territory, where a page leads to a few children and stops. But mesh density does not buy resilience. The densest site in the cohort still concentrates 39% of its betweenness in 1% of its pages. Connecting everything to a few central pages is not the same as distributing the load: a dense mesh wired through a small core is exactly as fragile as a sparse tree with the same core.

Content Quality at Scale

Across all twelve sites combined: 10,731 pages, 15.1 million tokens of body text, and roughly 136,000 internal editorial links between them.

Content metricSample mean (n=12)Notes
Pages per site894Median 995, range 441–1,001
Avg word count per page854Median 626 across sites
Avg token count per page~1,359Character-based estimate (length / 4)
Pages with thin content (under 200 words)26.5%Range 0%–66%
Pages with zero content~0.7%The “redirect / shell page” rate
Internal-link share of all links80% internalSites keep most links on-domain
Title tag coverage99.5%Near-universal
Meta description coverage86.6%One prestige site sits at 11%, dragging the mean

Two figures stand out. Thin content runs at 26.5% on average, roughly one page in four with fewer than 200 words: locale stubs, regulatory landing pages, single-claim product cards. On the most affected site, two-thirds of pages are thin. From an AI agent’s perspective these are weak signals, too short to anchor a confident summary but numerous enough to dilute a knowledge graph.

The meta-description figure hides a bimodal split. Most sites sit between 84% and 100% coverage, but one prestige site covers only 11% of its pages with a description. Title coverage is near-universal at 99.5%, so the omission is specifically the descriptive snippet that summarizes a page for both search engines and language models. On a site that already orphans most of its pages, the missing descriptions remove the one remaining out-of-band signal a crawler could have used.

What Every Finance F500 Site Shares

Where the healthcare cohort organized itself around the newsroom, the finance cohort organizes itself around jurisdiction. On several sites, the single largest URL section is a geographic or locale wrapper (/us, /global, /worldwide), and on those it absorbs the large majority of all pages. Some sites nest essentially their entire content tree under one locale path.

This single architectural choice cascades into the topology metrics above. When a site nests its entire content tree under a jurisdiction, the locale page becomes a mandatory bridge: every path to a product runs through it. That is precisely the betweenness concentration the fragility section measured. Jurisdiction silos also rarely cross-link: the regional trees do not point at each other, and none points at the global product catalog directly. The result is the topology we observed: a few load-bearing locale hubs, high orphan rates in the silos that nothing links back to, and a product story reachable only by first passing through a regulatory gateway.

The consumer-facing brands (the card, brokerage, and lending sites) escape some of this by organizing around products instead of jurisdictions; their dominant section absorbs a smaller share of pages. They have lower orphan rates and shallower paths. But they pay in dead-ends and thin content: product cards that receive links and lead nowhere.

What This Means for AI Search Readiness

For an AI agent that uses internal link structure to discover and rank content, three implications follow from this data.

1. Content can exist and still be invisible. This is the most consequential pattern here because it is silent: the pages render, the sitemap lists them, SEO audits pass, yet an agent following editorial links never reaches them. On the worst sites the majority of published content is dark. The fix is not more content; it is re-linking the archive into the live editorial graph, or accepting that for AI agents it does not exist.

2. Where a site blinds the tools, you cannot audit it at all. On two sites, both payment networks, invisibility was not even measurable: the anti-bot posture that protects a login flow also blinds the search engines and AI crawlers deciding whether the content is discoverable. For those, the first step is not a topology fix but making the inventory legible to the machines doing the reading.

3. The failure modes are independent, so one fix is never enough. Orphans (missing entry points) and dead-ends (circulation termini) fall on different sites, so closing one does nothing for the other, and neither touches the invisible tier a layer deeper. Each has to be measured and addressed on its own terms.

A note on prestige: the sites that hide the most content and concentrate the most betweenness are among the best-known institutions in the cohort. Brand authority and balance-sheet size do not buy a navigable topology. If anything, the global-jurisdiction architecture large institutions adopt works against it.

The fix is structural, and every piece of it is measurable: re-link invisible archives, spread betweenness off the few hubs, close dead-ends, rescue orphans into their communities, cross-link the jurisdiction silos, and make the sitemap itself reachable. This is the entire premise of the Digital MRI service.

Methodology

Twelve anonymized Fortune 500 financial-services sites, each run through the same pipeline: an HTTP-first adaptive crawl, main-content extraction (navigation, header, and footer links filtered out), dual-graph construction, and the five-lens topology analysis (Skeleton, Circulation, Organs, Health, Nervous System). This is a curated cohort, not a random sample, so the aggregate figures are exploratory benchmarks for F500 finance topology rather than point estimates for the sector. Reachability and invisibility use only the ten sites whose crawl captured a sitemap; topology, fragility, content, and the scorecard use all twelve, because those do not depend on the sitemap. One site is a large bank whose corporate domain is a thin holding site, so its consumer-banking property (a separate retail domain) is the one crawled here.

The dual-graph model. Each page is classified against two graphs from the same crawl: the full graph (navigation included) for reach, and the content graph (in-content editorial links only) for authority. A page is reachable via editorial links if it has an inbound content link, navigation-only if reachable in the full graph but not the content graph, and link-invisible if nothing links to it in either, surviving only in the sitemap. This refines the homepage-anchored model used in IR-001 (Healthcare, Q1 2026), which traversed from each homepage and so was sensitive to the chosen root; the dual-graph model is entry-point invariant. The two reports’ tiers are defined differently and their per-site rates are not directly comparable, but the shared finding holds in both: a measurable share of F500 content sits outside any link path an unaided agent can walk.

Disclaimers:

  • Methods. PageRank, Louvain community detection, and betweenness centrality over crawled page structure.
  • Robots & ethics. Site-level robots directives respected; disallowed pages never fetched. Publicly accessible page structure only. No content, metadata, or user data stored.
  • Anonymization. Codenames SF1-SF12; the codename-to-domain mapping is intentionally not published. A site-level data appendix is available on request.
  • Navigation exclusion. Any link target appearing on more than 80% of a site’s pages is treated as global navigation and dropped from the content graph; the full graph keeps it for reachability only.
  • Token counts. Character-based estimates (page length divided by four), not tokenizer output.
  • Betweenness. Reported as the share of total betweenness centrality carried by each site’s top 1% of pages.
  • Scope. Statistical patterns for educational purposes only; not advice about any specific site or company.