Jul 2026 Data Featured

Do Two AI Engines Cite the Same Web? No — 22% overlap

We ran one 100-query set through both Claude and Perplexity and compared every citation. They cited 793 domains between them — and agreed on 171. The overlap is 22%, and 16 of the 100 questions produced no shared source at all.

Queries
100
Union domains
793
Shared
171
Overlap
22%
Read time
6 min

Every finding on this site so far has come from a single engine. The long tail, the recurring head, the ~6 sources per answer — all Claude with web search. A fair question hangs over all of it: is any of this about the engine, or about the web? If a second engine cited the same domains, “who AI cites” would be a property of the internet. If it didn’t, then every citation number is engine-specific, and choosing which engine to optimize for is the first decision, not a footnote.

So we ran the experiment properly. One query set — the same 100 commercial and GEO/SEO questions — through two engines: Claude (Sonnet 4.6 + web search) and Perplexity (Sonar). Same questions, two models, runs a few days apart. Then we compared every cited domain, query by query.

Between them the two engines cited 793 distinct domains. They agreed on 171. That’s a 22% overlap — and on 16 of the 100 queries they shared no cited domain at all.

Two engines, two different webs

If the citation web were a property of the internet, the two domain sets would mostly coincide. They don’t. Claude cited 462 domains; Perplexity cited 502; only 171 sit in both sets. Put another way: 63% of the domains Claude cited, Perplexity never did — and 66% of Perplexity’s were invisible to Claude.

The per-query picture is starker than the totals suggest. For the average question, the two engines shared just 1.8 domains — against six or eight cited on each side. Their average per-query overlap is 15%. These aren’t two views of one ranked web; they’re two different retrieval systems reaching into two different corners of it and coming back with mostly different pages.

Where they do agree: the SEO-tool spine

The 171-domain intersection isn’t random. It’s dominated by the working toolkit of the SEO and content-marketing trade — the same names that formed Claude’s recurring head:

  • semrush.com — 12 of Claude’s answers, 15 of Perplexity’s. The single most-cited domain on both engines.
  • searchengineland.com, ahrefs.com, zapier.com, blog.hubspot.com, conductor.com — each cited several times by both.

If there’s such a thing as a cross-engine authority for this space, this is it: a couple dozen tools and trade publishers that both systems treat as reference material. For a GEO strategy, these are the highest-value targets — being cited here shows up no matter which engine your customer uses.

Where they split: social vs. editorial

The divergence has a shape, and it’s the most useful part of the whole comparison. The engines don’t just cite different domains at random — they lean toward different kinds of source:

  • Perplexity reaches for the social and community web. reddit.com in 89 of 100 answers (Claude: 0). youtube.com in 58 (Claude: 0). linkedin.com, g2.com, aioseo.com — forums, reviews, video, vendor pages. Perplexity’s answer is a synthesis of what people and vendors are saying.
  • Claude reaches for the editorial and publisher web. medium.com in 7 answers that Perplexity never cited; arxiv.org, clickrank.ai, shopify.com skewed heavily its way. Claude’s answer is built from articles and documentation.

Reddit is the clearest single fact in the dataset: the same 100 questions produced Reddit in 89 of Perplexity’s answers and 0 of Claude’s. That’s not a ranking difference — it’s a different theory of what a trustworthy source is.

What this means for GEO

  1. “Get cited by AI” is not one goal. It’s at least two, and they barely overlap. A page that wins on Perplexity (community proof, presence on Reddit and review sites) may be invisible to Claude, which wants editorial articles and docs — and vice versa. Decide which engine your audience uses before you optimize for “AI.”
  2. The shared spine is the safe bet. The ~171 domains both engines cite — led by the SEO tools — are where a citation pays off regardless of engine. If you have to pick one target, pick from the intersection.
  3. Single-engine numbers are single-engine numbers. Every “X% of AI citations” stat you read (including ours) describes one engine unless it says otherwise. A finding that doesn’t name its engine is describing a model, not the web — and the next model over disagrees 78% of the time.

Honest limits

  • Two engines, one basket. Claude Sonnet 4.6 and Perplexity Sonar, over one English-language, commercial/GEO/SEO-weighted 100-query set. A third engine would draw a third web; a different basket would move the specifics. The structure — low overlap, a small shared spine, engine-specific leans — is the durable read.
  • Not simultaneous. The two runs were days apart, so a sliver of the divergence is the live web changing underneath them, not the models. The gap is far too large (78%) for timing to explain, but it isn’t zero.
  • Reproducible. Both engines’ raw per-query citations are committed (raw-latest.json, raw-latest-perplexity.json); the overlap figures above are recomputed from them by scripts/build-cross-engine.mjs and checked in CI, so nothing here is hand-entered. Re-run it and you’ll get the same 22%.

The number worth carrying: two AI engines, asked the same hundred questions, cite the same source only 22% of the time. There is no single “AI citation web.” There’s Claude’s, and there’s Perplexity’s, and mostly they don’t touch.