
Research Delivered to Your Inbox
Get weekly, data-backed insights on AI visibility and search visibility optimization.
Most AI citations are rewrites, not quotes.
76% of AI citations are syntheses — the LLM uses your page as raw material and writes its own answer. Only 24% (593 of 2,422 analyzed sentences) are traceable to a specific passage.
When AI does quote you, it takes one sentence.
The typical quoted passage is 25 tokens — about 19 words — and 97% of all quoted passages across every surface stay under 200 tokens.
Verbatim quotes go to brand pages, never to listicles.
Owned brand pages get traced verbatim 84% of the time. Third-party listicles: 0%.
Google AIO is the only surface that quotes long.
Its median quoted chunk is 50 tokens — double the pooled median — and it occasionally pulls passages up to 500 tokens.
AI distills your content, it doesn't read it.
The 25-token ceiling is the design constraint most content strategies haven't been built around yet.
Abstract
For two years, every GEO playbook told the same story. Earn citations. Build authority. Win AI visibility. The implicit assumption behind that advice was consistent: A citation means the LLM is engaging with your content, lifting from it, attributing your ideas, and expressing your brand in its answer. Nobody tested what a citation actually looks like at the passage level. We did. We traced 2,422 AI-generated sentences back to their source pages and asked a simple question: how closely does what the AI wrote match what the cited page actually says? The finding changes how you should think about what a citation is worth.
How We Traced 2,422 AI-cited Sentences Back to Their Sources
Across 5 surfaces, we reverse-mapped 2,422 cited sentences to their source content.
We collected AI-generated responses across Google AI Overviews, ChatGPT, Perplexity, Claude, and Gemini.
For each cited URL, we fetched the source page, chunked it into overlapping segments of 25–500 tokens, and computed cosine similarity between each AI sentence and every chunk on the source page.
A "traceable citation" required cosine similarity ≥ 0.80. It is a metric used to measure how semantically related two pieces of text are. A score closer to 1 means the texts are nearly identical in meaning, while 0 means they're unrelated.
At 0.80, the AI sentence and source passage are saying substantively the same thing. Anything below that threshold was classified as synthesis, meaning the sentence was informed by the page, but couldn't be traced back to any single passage.
Result: 593 traceable sentences out of 2,422 total.
The other 76% we call the Synthesis Tax: Citations where your page shaped the AI's answer, but no specific passage was traced back to the source.
Finding 1: Three-Quarters of Every Citation Is Invisible Labor
76% of AI-cited sentences cannot be traced to a single source passage.
When an LLM cites your page, the most likely thing it's doing is absorbing your page, the argument, structure, and terminology and then writing its own sentence. Your content shaped the answer. Your brand, however, may not appear in it. Your ideas are present. Your words, however, are not.
This is the Synthesis Tax: the work you did to create the content, paid as a toll to the AI's answer, with no attribution to any specific passage.
The 24% that are traceable are cases where the AI stuck close to your wording, close enough that the semantic fingerprint survived. They're the exception, not the rule.
Finding 2: The Passages That Get Quoted Are Very Short
Median quoted chunk: 25 tokens (~19 words). 97% of all traceable citations are under 200 tokens.
| Surface | Traceable (N) | Typical quote length | 8 in 10 quotes are under... | % that stay under the paragraph |
|---|---|---|---|---|
| Google AIO | 41 | ~38 words | ~150 words | 83% |
| ChatGPT | 70 | ~19 words | ~75 words | 93% |
| Perplexity | 36 | ~19 words | ~38 words | 94% |
| Claude | 285 | ~19 words | ~19 words | ~100% |
| Gemini | 161 | ~19 words | ~38 words | 99% |
| All surfaces | 593 | ~19 words | ~38 words | 97% |
Note:
- N = number of traceable citations

What 25 tokens looks like in practice:
"Gainsight's health scoring assigns weighted scores across product usage, support tickets, and NPS to predict churn risk."
One sentence. One claim. That's what gets used.
The LLM isn't citing your 4,000-word guide because it needed the whole guide. It found one sentence or less and it used that.
Claude is the most extreme case: every traceable citation is under 200 tokens, with P80 at just 25 tokens. Gemini is nearly identical. Google AIO is the only surface pulling from longer, denser passages; its P95 is 500 tokens, meaning AIO occasionally surfaces material that every other model ignores.
Finding 3: Your Own Pages Get Quoted, Listicles Get Synthesized
When we enriched traceable citations with source-type data, a clear pattern emerged across all five surfaces.
| Source type | Times cited | Traceable sentences | Traceable rate |
|---|---|---|---|
| Owned brand pages | 179 | 150 | 84%* |
| Community (Reddit, forums) | 4 | 1 | 25% |
| Third-party listicles | 2 | 0 | 0%* |

Owned brand pages, content you control on your own domain, get quoted nearly verbatim 84% of the time. LLMs lift from owned pages with precision.
Third-party listicles appear in citation footnotes but generate zero directly traceable sentences. When an LLM cites a "10 best CRM tools" article, it synthesizes across the entire list rather than quoting a single line.
Citation events come in two fundamentally different flavors:
- Getting featured on a listicle builds mention presence. The LLM may name your brand in the context of a ranked list. But it is not quoting you.
- Getting your own page cited builds quote presence. When the LLM cites an owned page, there is an 84% chance it is pulling from a specific passage you wrote.
Finding 4: Brand Authority Doesn't Earn You Longer Quotes
We grouped traceable citations by brand tier, from Q1 category leaders to Q4 challengers, using domain authority as the divider. The question: do established brands earn richer, more direct quotes?
They don't.
| Brand tier | Traceable citations | P50 chunk (tokens) | P80 chunk (tokens) |
|---|---|---|---|
| Q1 leaders (DA ≥ 71) | 12 | 25 | 50 |
| Q3 mid-market (DA 31–50) | 39 | 25 | 50 |
| Q4 challengers (DA ≤ 30) | 7 | 25 | 25 |

The 25-token ceiling applies equally to a category leader and a Series A startup. Quote length is not a function of brand authority, it is a function of how LLMs process text.
(Note: Brand-tier coverage in our traceable dataset is 10.5%, 62 of 593 matched to a known brand, so these numbers are directional. The pattern is consistent: no tier advantage in how you get quoted, only in whether you get cited.)
ACTION ITEMS
Audit your most-cited pages for sentence density
Move your key claims to the first paragraph
Write 25-token facts
Separate your listicle strategy from your quote strategy
Accept the 76% as a structural baseline, not a failure
Get Your Personalized Action Plan
Receive a tailored action plan with clear, actionable recommendations you can start implementing right away.
ABOUT THE RESEARCHER

Head of SEO Research
“Pushkar leads SEO Research at VisibilityStack, driving the development of proprietary methodologies and frameworks that power our platform. His deep expertise in search algorithms and AI systems informs our technical approach. Pushkar has led SEO research initiatives at multiple technology companies, developing frameworks that have driven hundreds of millions in organic pipeline for B2B SaaS clients.”





