
5 min
Zach Chmael
In This Article
The founding GEO paper has two results tables that disagree. The one with error bars says keyword stuffing did nothing at all. Almost nobody has read it.
Updated
TL;DR
📄 The paper everyone calls "the Princeton study" was led by a researcher at IIT Delhi, with three Princeton co-authors and two independent researchers. The author order also flipped between the preprint and the published version.
📊 The famous 40% is a ceiling, not an average. The three best methods produced 30–40% relative improvement on one metric and 15–30% on the other.
⚠️ The paper reports the same experiment twice. Table 1 has no error bars; Table 6 does, and the numbers differ.
🔑 In the table with error bars, keyword stuffing scored 19.8 against a baseline of 19.8. The field's favorite finding, that keyword stuffing backfires, does not survive its own confidence intervals.
🎭 "Authoritative tone does nothing" is half a result. It moved nothing on word count and produced 23.1 against a 19.8 baseline on the subjective metric.
🧪 The real-world Perplexity test used 200 samples with sources uploaded as files, and Cite Sources scored 19.0 against a 24.7 baseline on subjective impression. It made things worse.
The Most-Quoted Study in AI Search, Quoted Correctly
There is one paper under nearly every claim made about AI search. It has been summarized thousands of times, and I have yet to find a summary that survives contact with the appendix.
I went back to it last week to check figures for two other articles. I expected to confirm what I already believed and move on. Instead I found that the paper contains two results tables covering the same experiment, that they disagree, that the one almost nobody cites has error bars, and that once you read them, at least two of the field's most-repeated takeaways stop being supported. Including two I had published myself a few days earlier.
This is that reading. Every figure below comes from the published version, not from anyone's summary of it.
Who actually wrote the GEO paper?
Six researchers across three affiliations, led by someone who was not at Princeton. The author block lists Pranjal Aggarwal at the Indian Institute of Technology Delhi, Vishvak Murahari at Princeton, Tanmay Rajpurohit and Ashwin Kalyan as independent researchers in Seattle, and Karthik Narasimhan and Ameet Deshpande at Princeton.
Aggarwal and Murahari are marked as equal contributors. And the order between them flipped: the project's own citation block lists Murahari first, matching the original 2023 preprint, while the KDD proceedings entry lists Aggarwal first. Which name leads depends on which version you are looking at.
"The Princeton study" is the shorthand the entire industry uses. It erases the lead author's institution and flattens a six-person, multi-institution collaboration into one university's brand.
Before this reads as pedantry, consider how far the sloppiness travels.
Two later arXiv papers citing this work attribute it to authors who do not appear on it at all: one credits "Aggarwal, A., Kanakia, A., Sharma, A., & Chang, M. W.," and another lists "Muralidhar N. Aggarwal, P. and A. Nagar." Those are academic preprints, not marketing blogs. When a citation degrades this badly this fast, the number attached to it deserves checking too.
What did the study actually measure?
A simulated two-stage engine, not a live product. Google returned the top five sources for a query, then GPT-3.5-turbo wrote a cited answer using only those sources. Five responses were sampled at temperature 0.7 across five random seeds.
The benchmark, GEO-bench, holds 10,000 queries split 8,000 / 1,000 / 1,000 across nine datasets and 25 domains, preserving a real-world mix of 80% informational and 10% each transactional and navigational. Sources include MS MARCO, ORCAS, Natural Questions, ELI5, and queries generated by GPT-4.
Two metrics carry every claim in the paper. Position-Adjusted Word Count measures how much of the answer traces to your source, weighted by an exponential decay on position. Subjective Impression is a seven-part score covering relevance, influence, uniqueness, diversity, position, count, and likelihood of a click, computed by GPT-3.5 using the G-Eval method.
They frequently disagree with each other, which turns out to matter enormously.
Three constraints worth holding onto. The engine was GPT-3.5-turbo. Only five sources entered the context window, for cost reasons. And the authors state plainly in their limitations that they did not evaluate how any of these methods affect search rankings.
Is the 40% figure an average?
No. It is a maximum, and the paper says so in its own abstract: GEO can boost visibility up to 40%. The three strongest methods, Cite Sources, Quotation Addition, and Statistics Addition, produced a relative improvement of 30–40% on Position-Adjusted Word Count and 15–30% on Subjective Impression. The single best method improved on baseline by 41% on the first metric and 28% on the second.
That is a real result. It is not "GEO increases visibility by 40%," which is how it usually arrives.
Why do the paper's two results tables disagree?
Because Table 1 reports single values and Table 6, in the appendix, reports the same experiment with standard deviations across five seeds. The numbers are not the same, and the gap is not cosmetic.
Method | Table 1 (PAWC) | Table 6 (PAWC, ±σ) |
|---|---|---|
No optimization | 19.3 | 19.8 (±0.6) |
Keyword Stuffing | 17.7 | 19.8 (±0.5) |
Unique Words | 20.5 | 20.7 (±0.5) |
Easy-to-Understand | 22.0 | 21.5 (±0.6) |
Authoritative | 21.3 | 21.1 (±0.8) |
Technical Terms | 22.7 | 22.5 (±0.6) |
Fluency Optimization | 24.7 | 24.4 (±0.8) |
Cite Sources | 24.6 | 25.3 (±0.6) |
Quotation Addition | 27.2 | 27.1 (±0.6) |
Statistics Addition | 25.2 | 25.5 (±1.2) |
For the top three methods the two tables broadly agree, and their gains sit comfortably outside the error bars. Quotation Addition at 27.1 (±0.6) against a 19.8 (±0.6) baseline is a real effect by any reading. That part of the paper holds.
The bottom of the table is where the field's summaries fall apart.
Did keyword stuffing actually hurt?
Not according to the table with error bars. Keyword Stuffing scored 19.8 (±0.5). The baseline scored 19.8 (±0.6). Identical, well inside one standard deviation.
The claim that keyword stuffing actively backfires in generative engines rests entirely on Table 1's pair of 17.7 against 19.3. It is repeated constantly, usually as the punchy proof that old SEO habits are now counterproductive. I published it myself twice in the past week.
What the study supports is weaker and more useful: keyword stuffing does nothing. The paper's own prose says exactly that, describing methods that "offer little to no improvement." Little to no improvement is not the same as harm, and the difference matters if you are deciding whether to spend a sprint removing keyword density from old pages. On this evidence, that sprint buys you nothing in either direction.
Does sounding authoritative really do nothing?
Only on one of the two metrics, and the field consistently reports the half that is more quotable.
On Position-Adjusted Word Count, the Authoritative method scored 21.1 (±0.8) against a 19.8 (±0.6) baseline. Marginal, roughly within noise, and the basis for the widely-repeated line that generative engines are already resistant to tonal manipulation. The paper draws that conclusion itself.
On Subjective Impression, the same method averaged 23.1 (±0.7) against a 19.8 (±0.9) baseline, one of the larger gains in the whole table and ahead of Fluency Optimization, Technical Terms, and Easy-to-Understand.
So a more persuasive, confident register did not get your source quoted at greater length or in a better position. It did make the resulting citation read as more relevant, more influential, and more clickable to the evaluator. Both things are true, and "authoritative tone does nothing" reports one of them.
I am not arguing you should go write in a more authoritative voice. I am pointing out that the summary everyone repeats was produced by picking the metric that made the cleaner story.
Where does the "+115% for page five" number come from?
Section 5.2, a deliberately separate experiment in which every source was optimized at once rather than one per query. The authors wanted to model what happens when GEO becomes universal.
SERP rank | Cite Sources | Quotation Addition | Statistics Addition |
|---|---|---|---|
Rank 1 | −30.3% | −22.9% | −20.6% |
Rank 2 | +2.5% | −7.0% | −3.9% |
Rank 3 | +20.4% | +3.5% | +8.1% |
Rank 4 | +15.5% | +25.1% | +10.0% |
Rank 5 | +115.1% | +99.7% | +97.9% |
The paper reads this optimistically, as evidence that generative engines could democratize a space where backlinks and domain authority favor incumbents. That is a fair reading of the data.
But it describes redistribution under saturation, not individual uplift. In the same condition that hands the fifth-ranked page 115%, the top-ranked page loses 30.3%. Quoting the gain without the loss converts a finding about a zero-sum shift into a promise about your own page, which is a different claim with a different expected value. If you are ranked fifth today, this helps you most in a world where competitors also do it, and least in a world where only you do.
How real-world was the real-world test?
Less than the framing suggests. The section titled "GEO in the wild" evaluated methods on Perplexity.ai using 200 samples from the test split, not the full benchmark. And because Perplexity does not let a user specify source URLs, the researchers uploaded the source text as files and constrained answers to those uploads.
That is a reasonable workaround and it is not the same thing as optimizing a live web page and watching a production engine find it. No crawling, no indexing, no ranking, no competition with the rest of the web. The paper is transparent about all of this. Secondary coverage rarely is.
What did the Perplexity results actually show?
Something considerably messier than the summaries suggest, including two results that point the opposite way from the headline findings.
Method | PAWC (baseline 24.1) | Subjective Impression (baseline 24.7) |
|---|---|---|
Keyword Stuffing | 21.9 | 28.1 |
Cite Sources | 26.8 | 19.0 |
Quotation Addition | 29.1 | 32.1 |
Statistics Addition | 26.2 | 33.9 |
Authoritative | 25.9 | 30.6 |
Read the first two rows against everything the field says about this paper.
Keyword stuffing improved subjective impression, 28.1 against a 24.7 baseline, while lowering word count. The technique everyone cites this study to condemn produced one of the better subjective scores on the only commercially deployed engine tested.
Cite Sources made subjective impression substantially worse, 19.0 against a 24.7 baseline. Adding citations, the single most universally recommended tactic in generative engine optimization, dropped the score by more than five points on a real engine. I have never once seen this reported.
Quotation Addition and Statistics Addition held up well on both metrics, which is the sturdiest finding here and the one worth acting on. But a study whose real-engine test shows its flagship recommendation degrading one of its two metrics is a study that deserves more hedging than it gets.
So what does the paper actually support?
Four claims, stated at the strength the evidence justifies.
Adding quotations and statistics from credible sources improves visibility. This is the strongest result in the paper. It holds across both metrics, in both results tables, outside the error bars, and on the deployed engine. If you act on one thing, act on this.
Clear writing is a visibility lever. Fluency Optimization reached 24.4 (±0.8) while adding no new information, and the best pairing in the combination analysis was Fluency Optimization with Statistics Addition at 35.8%.
Effects vary by domain. Cite Sources performed best on factual and legal queries; Quotation Addition on people, society, and history. Blanket tactics are the wrong unit.
Optimization redistributes attention rather than creating it. The rank-1 penalty under universal adoption is the paper's own evidence for this, and C-SEO Bench at NeurIPS 2025 reached the same structural conclusion from an independent benchmark: as adoption rises, marginal gains compress.
What it does not support: that keyword stuffing harms you, that authoritative tone is inert, that any single page should expect a 115% gain, or that citations reliably improve how your source is perceived on a live engine.
What did we publish that needs correcting?
Two articles on this site, both from the past week, contain the claim that keyword stuffing scored 17.7 against a 19.3 baseline and therefore performed worse than doing nothing. That figure is real and it comes from Table 1. It is also contradicted by Table 6, where the same comparison is 19.8 against 19.8.
The accurate claim is that keyword stuffing produced no measurable benefit. Both pieces have been updated. One of them also reported that authoritative tone produced no significant improvement without noting that this holds only for Position-Adjusted Word Count.
I would rather publish this than quietly edit the pages. A company whose position is measurement honesty does not get to correct itself silently, and a correction log is cheaper than the credibility it protects.
Why does any of this matter for your content?
Because the field is building strategy on a summary of a summary, and the errors compound in a predictable direction: toward tactics that are easy to sell.
The misquoted version of this paper says AI visibility is a solved technical problem with a 40% payoff and a list of moves. The actual paper says three content changes reliably help, several do nothing, one flagship tactic degraded a metric on the only live engine tested, the gains redistribute rather than accumulate, and the whole thing ran on a model that is now several generations old.
The second version is less marketable and more useful. It tells you to add real statistics and real quotations from credible sources, to write clearly, to stop expecting a fixed return, and to measure your own results because the published ones may not transfer.

How does Trovance handle this?
By treating published research as something to verify rather than cite. Every claim that reaches a Trovance recommendation traces to a primary source we have read, with its experimental conditions attached, because a recommendation built on a misquoted study is worse than no recommendation.
That discipline is also why the product measures your own answers rather than promising you someone else's numbers. The GEO results ran on GPT-3.5-turbo with five sources in 2024. Your buyers are asking different questions of different engines today, and the only reliable evidence about your visibility is the evidence you collect about your visibility.
Trovance does not guarantee citations. What it does is preserve the before-state, repeat the runs enough times to distinguish signal from variance, and hold the diagnosis against what the evidence actually supports.
Humans keep the consequential calls: whether a claim is true, whether the proof is sufficient, and what deserves to ship.
If you want to see what the engines currently say about your company, start with your own market questions and read the raw answers before you change anything.
FAQs
Is the GEO study a Princeton study?
Partly. Three of six authors were at Princeton, but lead author Pranjal Aggarwal was at IIT Delhi, and two co-authors were independent researchers. Calling it "the Princeton study" erases the lead author's institution. The accurate reference is Aggarwal et al., KDD 2024.
Does the GEO study prove AI visibility increases 40%?
It found up to 40%, which is a ceiling rather than an average. The three strongest methods produced 30–40% relative improvement on Position-Adjusted Word Count and 15–30% on Subjective Impression. The single best method reached 41% on the first metric.
Did the study find keyword stuffing hurts AI visibility?
Not in the table reporting error bars. There, keyword stuffing scored 19.8 against a baseline of 19.8, meaning no measurable difference. The widely-cited 17.7 versus 19.3 comparison comes from a separate table without confidence intervals. The supported claim is that it does nothing.
Where does the 115% figure for lower-ranked pages come from?
Section 5.2, a separate experiment in which all sources were optimized simultaneously to model universal adoption. In that same condition, the top-ranked page lost 30.3% of its visibility. The figure describes redistribution under saturation, not the gain an individual page should expect.
Was the GEO study tested on real AI search engines?
Partially. The main experiments used a simulated engine running GPT-3.5-turbo over five Google results. A smaller test on Perplexity.ai used 200 samples with source text uploaded as files, because Perplexity does not allow users to specify source URLs.
Should I add citations to my content based on this research?
Adding quotations and statistics has the strongest support across every table in the paper. Citations are more mixed: strong in the simulated engine, but on Perplexity the Cite Sources method scored 19.0 against a 24.7 baseline on subjective impression, making it worse.
Is the GEO research still current?
Its core findings about statistics and quotations remain well supported, but the experiments ran on GPT-3.5-turbo with only five sources in context. The authors noted in their limitations that methods may need to adapt as engines evolve, and C-SEO Bench has since found gains compress as adoption rises.
Zach Chmael is co-founder of Trovance, where he works on the evidence and production layer for the agentic web. He writes about what the research on AI citation actually supports, including the parts that undercut the tactics.
Trovance shows how AI systems explain, compare, and recommend your company, identifies the evidence keeping you off the right shortlists, and turns those gaps into proof-backed content your team can publish. See how AI sees your brand.
