Do Two AI Engines Cite the Same Web? No — 22% overlap
We ran one 100-query set through both Claude and Perplexity and compared every citation. They cited 793 domains between them — and agreed on 171. The overlap is 22%, and 16 of the 100 questions produced no shared source at all.
- Queries
- 100
- Union domains
- 793
- Shared
- 171
- Overlap
- 22%
- Read time
- 6 min
Every finding on this site so far has come from a single engine. The long tail, the recurring head, the ~6 sources per answer — all Claude with web search. A fair question hangs over all of it: is any of this about the engine, or about the web? If a second engine cited the same domains, “who AI cites” would be a property of the internet. If it didn’t, then every citation number is engine-specific, and choosing which engine to optimize for is the first decision, not a footnote.
So we ran the experiment properly. One query set — the same 100 commercial and GEO/SEO questions — through two engines: Claude (Sonnet 4.6 + web search) and Perplexity (Sonar). Same questions, two models, runs a few days apart. Then we compared every cited domain, query by query.
Between them the two engines cited 793 distinct domains. They agreed on 171. That’s a 22% overlap — and on 16 of the 100 queries they shared no cited domain at all.
Two engines, two different webs
If the citation web were a property of the internet, the two domain sets would mostly coincide. They don’t. Claude cited 462 domains; Perplexity cited 502; only 171 sit in both sets. Put another way: 63% of the domains Claude cited, Perplexity never did — and 66% of Perplexity’s were invisible to Claude.
The per-query picture is starker than the totals suggest. For the average question, the two engines shared just 1.8 domains — against six or eight cited on each side. Their average per-query overlap is 15%. These aren’t two views of one ranked web; they’re two different retrieval systems reaching into two different corners of it and coming back with mostly different pages.
Where they do agree: the SEO-tool spine
The 171-domain intersection isn’t random. It’s dominated by the working toolkit of the SEO and content-marketing trade — the same names that formed Claude’s recurring head:
- semrush.com — 12 of Claude’s answers, 15 of Perplexity’s. The single most-cited domain on both engines.
- searchengineland.com, ahrefs.com, zapier.com, blog.hubspot.com, conductor.com — each cited several times by both.
If there’s such a thing as a cross-engine authority for this space, this is it: a couple dozen tools and trade publishers that both systems treat as reference material. For a GEO strategy, these are the highest-value targets — being cited here shows up no matter which engine your customer uses.
Where they split: social vs. editorial
The divergence has a shape, and it’s the most useful part of the whole comparison. The engines don’t just cite different domains at random — they lean toward different kinds of source:
- Perplexity reaches for the social and community web.
reddit.comin 89 of 100 answers (Claude: 0).youtube.comin 58 (Claude: 0).linkedin.com,g2.com,aioseo.com— forums, reviews, video, vendor pages. Perplexity’s answer is a synthesis of what people and vendors are saying. - Claude reaches for the editorial and publisher web.
medium.comin 7 answers that Perplexity never cited;arxiv.org,clickrank.ai,shopify.comskewed heavily its way. Claude’s answer is built from articles and documentation.
Reddit is the clearest single fact in the dataset: the same 100 questions produced Reddit in 89 of Perplexity’s answers and 0 of Claude’s. That’s not a ranking difference — it’s a different theory of what a trustworthy source is.
What this means for GEO
- “Get cited by AI” is not one goal. It’s at least two, and they barely overlap. A page that wins on Perplexity (community proof, presence on Reddit and review sites) may be invisible to Claude, which wants editorial articles and docs — and vice versa. Decide which engine your audience uses before you optimize for “AI.”
- The shared spine is the safe bet. The ~171 domains both engines cite — led by the SEO tools — are where a citation pays off regardless of engine. If you have to pick one target, pick from the intersection.
- Single-engine numbers are single-engine numbers. Every “X% of AI citations” stat you read (including ours) describes one engine unless it says otherwise. A finding that doesn’t name its engine is describing a model, not the web — and the next model over disagrees 78% of the time.
Honest limits
- Two engines, one basket. Claude Sonnet 4.6 and Perplexity Sonar, over one English-language, commercial/GEO/SEO-weighted 100-query set. A third engine would draw a third web; a different basket would move the specifics. The structure — low overlap, a small shared spine, engine-specific leans — is the durable read.
- Not simultaneous. The two runs were days apart, so a sliver of the divergence is the live web changing underneath them, not the models. The gap is far too large (78%) for timing to explain, but it isn’t zero.
- Reproducible. Both engines’ raw per-query citations are committed (
raw-latest.json,raw-latest-perplexity.json); the overlap figures above are recomputed from them byscripts/build-cross-engine.mjsand checked in CI, so nothing here is hand-entered. Re-run it and you’ll get the same 22%.
The number worth carrying: two AI engines, asked the same hundred questions, cite the same source only 22% of the time. There is no single “AI citation web.” There’s Claude’s, and there’s Perplexity’s, and mostly they don’t touch.