GEO

AI Doesn't Quote You, It Rewrites You: 76% of Citations Prove It

Researchers:
Pushkar Sinha
Peer Reviewers:
Ameet Mehta
Published date:Jul 17, 2026
4 min read
AI Doesn't Quote You, It Rewrites You: 76% of Citations Prove It
01

Most AI citations are rewrites, not quotes.

76% of AI citations are syntheses — the LLM uses your page as raw material and writes its own answer. Only 24% (593 of 2,422 analyzed sentences) are traceable to a specific passage.

02

When AI does quote you, it takes one sentence.

The typical quoted passage is 25 tokens — about 19 words — and 97% of all quoted passages across every surface stay under 200 tokens.

03

Verbatim quotes go to brand pages, never to listicles.

Owned brand pages get traced verbatim 84% of the time. Third-party listicles: 0%.

04

Google AIO is the only surface that quotes long.

Its median quoted chunk is 50 tokens — double the pooled median — and it occasionally pulls passages up to 500 tokens.

05

AI distills your content, it doesn't read it.

The 25-token ceiling is the design constraint most content strategies haven't been built around yet.

Abstract

For two years, every GEO playbook told the same story. Earn citations. Build authority. Win AI visibility. The implicit assumption behind that advice was consistent: A citation means the LLM is engaging with your content, lifting from it, attributing your ideas, and expressing your brand in its answer. Nobody tested what a citation actually looks like at the passage level. We did. We traced 2,422 AI-generated sentences back to their source pages and asked a simple question: how closely does what the AI wrote match what the cited page actually says? The finding changes how you should think about what a citation is worth.

How We Traced 2,422 AI-cited Sentences Back to Their Sources

Across 5 surfaces, we reverse-mapped 2,422 cited sentences to their source content.

We collected AI-generated responses across Google AI Overviews, ChatGPT, Perplexity, Claude, and Gemini.

For each cited URL, we fetched the source page, chunked it into overlapping segments of 25–500 tokens, and computed cosine similarity between each AI sentence and every chunk on the source page.

A "traceable citation" required cosine similarity ≥ 0.80. It is a metric used to measure how semantically related two pieces of text are. A score closer to 1 means the texts are nearly identical in meaning, while 0 means they're unrelated.

At 0.80, the AI sentence and source passage are saying substantively the same thing. Anything below that threshold was classified as synthesis, meaning the sentence was informed by the page, but couldn't be traced back to any single passage.

Result: 593 traceable sentences out of 2,422 total.

The other 76% we call the Synthesis Tax: Citations where your page shaped the AI's answer, but no specific passage was traced back to the source.

Finding 1: Three-Quarters of Every Citation Is Invisible Labor

76% of AI-cited sentences cannot be traced to a single source passage.

When an LLM cites your page, the most likely thing it's doing is absorbing your page, the argument, structure, and terminology and then writing its own sentence. Your content shaped the answer. Your brand, however, may not appear in it. Your ideas are present. Your words, however, are not.

This is the Synthesis Tax: the work you did to create the content, paid as a toll to the AI's answer, with no attribution to any specific passage.

The 24% that are traceable are cases where the AI stuck close to your wording, close enough that the semantic fingerprint survived. They're the exception, not the rule.

Business implication: The number to internalize: For every 4 times your content is cited, 3 are synthesis events. Only 1 is a traceable quote. If you measure citation volume as a proxy for how your brand is being expressed in AI answers, you are measuring the right thing for only 25% of your citations.

Finding 2: The Passages That Get Quoted Are Very Short

Median quoted chunk: 25 tokens (~19 words). 97% of all traceable citations are under 200 tokens.

SurfaceTraceable (N)Typical quote length8 in 10 quotes are under...% that stay under the paragraph
Google AIO41~38 words~150 words83%
ChatGPT70~19 words~75 words93%
Perplexity36~19 words~38 words94%
Claude285~19 words~19 words~100%
Gemini161~19 words~38 words99%
All surfaces593~19 words~38 words97%

Note:

  • N = number of traceable citations
AI Models and the token lengths they cite

What 25 tokens looks like in practice:

"Gainsight's health scoring assigns weighted scores across product usage, support tickets, and NPS to predict churn risk."

One sentence. One claim. That's what gets used.

The LLM isn't citing your 4,000-word guide because it needed the whole guide. It found one sentence or less and it used that.

Claude is the most extreme case: every traceable citation is under 200 tokens, with P80 at just 25 tokens. Gemini is nearly identical. Google AIO is the only surface pulling from longer, denser passages; its P95 is 500 tokens, meaning AIO occasionally surfaces material that every other model ignores.

Business implication: Implication for content structure: Long-form content still earns citations. But the citation value lives in individual sentences, not paragraphs. If your best factual claims are buried in paragraph four, the LLM may never reach them. Front-load the fact.

Finding 3: Your Own Pages Get Quoted, Listicles Get Synthesized

When we enriched traceable citations with source-type data, a clear pattern emerged across all five surfaces.

Source typeTimes citedTraceable sentencesTraceable rate
Owned brand pages17915084%*
Community (Reddit, forums) 4125%
Third-party listicles200%*
Traceable Rate Comparison

Owned brand pages, content you control on your own domain, get quoted nearly verbatim 84% of the time. LLMs lift from owned pages with precision.

Third-party listicles appear in citation footnotes but generate zero directly traceable sentences. When an LLM cites a "10 best CRM tools" article, it synthesizes across the entire list rather than quoting a single line.

Citation events come in two fundamentally different flavors:

  • Getting featured on a listicle builds mention presence. The LLM may name your brand in the context of a ranked list. But it is not quoting you.
  • Getting your own page cited builds quote presence. When the LLM cites an owned page, there is an 84% chance it is pulling from a specific passage you wrote.
Business implication: Listicle features and owned-page citations are different citation types with different downstream effects on brand expression. Measuring them together masks the distinction that matters most.

Finding 4: Brand Authority Doesn't Earn You Longer Quotes

We grouped traceable citations by brand tier, from Q1 category leaders to Q4 challengers, using domain authority as the divider. The question: do established brands earn richer, more direct quotes?

They don't.

Brand tierTraceable citationsP50 chunk (tokens)P80 chunk (tokens)
Q1 leaders (DA ≥ 71)122550
Q3 mid-market (DA 31–50)392550
Q4 challengers (DA ≤ 30)72525
Token Length Distribution by Brand Tier

The 25-token ceiling applies equally to a category leader and a Series A startup. Quote length is not a function of brand authority, it is a function of how LLMs process text.

(Note: Brand-tier coverage in our traceable dataset is 10.5%, 62 of 593 matched to a known brand, so these numbers are directional. The pattern is consistent: no tier advantage in how you get quoted, only in whether you get cited.)

Business implication: The practical point: You are not competing against Salesforce for a longer, richer quote. You are competing for the same 25-token slot. That field is level.
Action Items Icon

ACTION ITEMS

01

Audit your most-cited pages for sentence density

02

Move your key claims to the first paragraph

03

Write 25-token facts

04

Separate your listicle strategy from your quote strategy

05

Accept the 76% as a structural baseline, not a failure

Get Your Personalized Action Plan

Receive a tailored action plan with clear, actionable recommendations you can start implementing right away.

ABOUT THE RESEARCHER

Pushkar Sinha
Pushkar Sinha

Head of SEO Research

Pushkar leads SEO Research at VisibilityStack, driving the development of proprietary methodologies and frameworks that power our platform. His deep expertise in search algorithms and AI systems informs our technical approach. Pushkar has led SEO research initiatives at multiple technology companies, developing frameworks that have driven hundreds of millions in organic pipeline for B2B SaaS clients.

Methodology & History

SHARE THIS RESEARCH

Research Delivered to Your Inbox

Get weekly, data-backed insights on AI visibility and search visibility optimization.

Newsletter study mockup