Jul 2026 Data Featured

The Long Tail of LLM Citations: 85% of Domains Appear Once

We ran 100 commercial-intent queries through one engine — Claude with web search — and logged every citation. Of 462 unique domains, 395 showed up for a single query. The concentration is extreme — and it held when we scaled the sample up.

Traces
100
Citations
610
Domains
462
Cited once
85%
Read time
5 min
85%of cited domains appear for just one query
Embed this chart
↓ Download the share card (PNG)

When you run a GEO Trace and watch a dozen domains scroll past, the instinct is “a dozen sources are competing for this query.” Run a hundred queries and the picture inverts: almost none of those domains compete for anything. They show up once and are never seen again.

We ran a single-engine panel — 100 commercial and GEO/SEO queries through the exact call GEO Trace uses (Claude Sonnet 4.6 with web search), logging every citation in every answer. Here is what the citation distribution actually looks like.

Of 462 unique domains cited across the run, 395 (85%) appeared for only one query. Just ten domains were cited across five or more of the hundred. One — Semrush — reached twelve.

What we measured

One hundred queries. Every one returned a non-empty citation set — 610 citations in total, 6.1 per answer on average. Those 610 citations resolved to 462 unique domains, distributed by how many of the 100 queries cited each one:

  • 395 domains (85%) — cited for exactly 1 query
  • 44 domains (10%) — cited for 2
  • 10 domains — cited for 3
  • 3 domains — cited for 4
  • 9 domains — cited for 5–7
  • 1 domain — cited for 12 (Semrush)

There is no “always-cited” source. The most-recurring domain — semrush.com — appeared in just 12 of 100 answers; medium.com next at 7; then a cluster of publishers and tools (ahrefs.com, arxiv.org, blog.hubspot.com, neilpatel.com, searchengineland.com) at 6. That small group is the entire “head.” Everything below it is tail.

We scaled the sample — and caught our own number

An earlier version of this run used just 30 queries and reported 93% cited-once. Scaling to 100 queries, the figure settled at 85%. That drop is the honest, expected consequence of a bigger sample: with more queries, more domains get a second chance to reappear, so the once-only share compresses. The finding — an overwhelming long tail — held; the exact percentage was sample-dependent, and we’d rather show you the replicated 85% than the flashier 30-query 93%. (This is why we’re standing up a monthly panel: one run is a snapshot, not a constant.)

Why the distribution looks like this

For a live-retrieval engine, each query triggers a fresh web search and a fresh candidate set. The model cites what is most relevant to that question — and relevance is intensely query-specific. “Best CRM for startups” and “how to reduce server response time for SEO” simply do not share source pages, so their citations don’t overlap. A small number of wide-coverage publishers and tools (Semrush, Medium, the SEO trades) straddle several queries; the rest of the cited universe is a different set every time.

One honest scoping note: this is one engine. For a very different picture from the same 100 queries, see what Perplexity cites — its head is Reddit, not SEO tools — and the head-to-head.

What this means for GEO

1. Recurrence beats appearance. Being cited once, for one query, is the median outcome — it tells you almost nothing. A domain that recurs across several queries in your space (the way semrush.com does here) is a structural authority worth studying. A single citation is noise until it repeats.

2. Most of the cited universe is reachable. If the competition for a given query is a handful of query-specific pages rather than a fixed canon, then the tail is winnable. You’re not trying to displace Wikipedia; you’re trying to be the most relevant page for one specific question.

3. Read the tail for intelligence. The 395 single-query domains are where specialist sources hide. For competitive research, that long-tail column is the most valuable part of a trace, not the least.

Honest limits

  • N = 100, one engine, one basket. This is a pilot panel, not a census: English-language, commercial- and GEO/SEO-weighted queries. The numbers describe this run. A different query set would move them — probably not the 80%+ headline, but the specifics, yes.
  • Single engine. Everything above is one model’s behavior; the cross-engine question is answered separately.
  • Reproducible. The raw per-query citations behind every number live alongside this essay in the repo (src/data/panel/raw-latest.json), produced by scripts/run-panel.mjs. Re-run it and you’ll get your own distribution.

The number worth carrying: across a hundred queries, 85% of the domains a single engine cited appeared exactly once. Citation isn’t a fixed leaderboard you climb. It’s re-decided, query by query.