TL;DR
🧪 Retested at scale, the rewrite playbook mostly evaporates: only 3 of 54 unilateral conditions in C-SEO Bench produced a statistically significant citation-rank gain.
📊 What did hold up is evidence: the original GEO paper improved on baseline by up to 41% by adding statistics, quotations and cited sources, measured as visibility inside a generated answer.
🕸️ Fix retrieval before phrasing. Cloudflare's agent-readiness study covers the 200,000 most visited domains, and several major AI crawlers read raw HTML without running your JavaScript.
🔗 A lot of the decision happens off your site: 57% of AI citations point to sources brands do not control.
📐 Size your evidence to the unit that competes: retrieved passages in one 2026 multi-hop chemistry retrieval benchmark averaged 188 tokens with a 56.26-token standard deviation.
🎯 Measure appearance rates across repeated runs. A 2026 variance-components study found run-to-run noise large enough to hide real effects in small samples.
Most of what circulates as advice for getting cited by AI traces back to one 2023 paper, published at KDD in 2024, and a rewrite checklist that later benchmarks could not reproduce. When ten conversational-SEO methods were retested against a shared benchmark, only 3 of 54 unilateral conditions produced a statistically significant citation-rank gain. The properties that do survive independent testing are duller than the checklist: pages an engine can fetch, evidence it can quote, and attribution it can resolve. This piece separates the two.
The short answer to what gets cited: a passage carrying a specific, sourced claim, on a page the engine can retrieve, from a source the engine already trusts. The original generative engine optimization paper improved on baseline by up to 41% using statistics, quotations and cited sources, and the published version is careful that its metric is position-adjusted visibility inside a generated answer. Phrasing tricks stacked on top of that mostly do nothing.
These are commercial stakes. G2 found AI chatbots are now the single largest influence on B2B shortlists, and 51% of software buyers now begin research inside an AI chatbot. Getting the causal story wrong here does not cost you a ranking. It costs you a quarter of work on rewrites that testing says are inert.
What does the strongest citation evidence actually say?
It says evidence-carrying text wins, and it says so under conditions worth naming out loud. The study ran nine optimization methods over GEO-bench, a query set built for the purpose, and reported the largest gains for adding statistics, quotations and cited sources. The KDD proceedings entry and the project's own citation block both make the scope explicit.
What it measured matters more than the headline percentage. The visibility metric is position-adjusted word count inside a generated answer. Analytics links and click counts play no part in it. A gain reported as "+41% visibility" means more of your text surfaced inside a controlled answer, which is a real signal and a different one from being named or shortlisted.
The same results carry a negative that gets dropped in the retellings. Keyword stuffing, the tactic that defined a decade of optimization, scored below the untouched baseline in the paper's own comparison. Any checklist quoting the positive findings while omitting that one is quoting selectively.
Why do most rewrite tactics fail when someone retests them?
Because the effect that survives lives in the evidence itself, and the wording wrapped around it carries almost nothing. C-SEO Bench took ten published conversational-SEO methods and tested them across two tasks and six domains, which is a wider surface than the original work covered. Almost nothing held: 3 significant conditions out of 54 is close to what you would expect from chance alone at a conventional threshold.
The benchmark's second result is the one that should reset planning. It also measured what happens when every document in the pool applies the same method, which is the realistic case in any category where competitors read the same blog posts. Advantages that looked real in isolation shrink when everyone adopts them, and the dataset paper points instead to traditional retrieval-side improvements as the durable path.
Read practically, this is permission to stop doing a category of work. Rewriting a page in a more conversational register, or restating the same claims in question form, are not interventions that testing supports. If the rewrite does not add a fact, a source, or a retrievable answer, treat its expected effect as zero until you measure otherwise.
Which content properties actually correlate with citation?
Three: pages the engine can fetch, evidence it can quote, and attribution it can resolve. Fix them in that order, because evidence on a page an engine cannot fetch is evidence nobody reads.
1. Pages the engine can fetch
Several major AI crawlers do not execute JavaScript. Vercel ran with MERJ a study of crawler behavior showing that the agents it measured read raw HTML, so a client-rendered page arrives close to empty. Google is the documented exception, because its indexing pipeline applies the same Web Rendering Service capabilities its search index has always used. Cloudflare surveyed the 200,000 most visited domains for agent readiness, which is the scale at which this question is now being asked.
2. Evidence a model can quote
Retrieval systems assemble answers out of passages, so the practical question is which passages get pulled. An analysis of 548,534 pages maps which page traits correlate with being pulled into answers, and a 2026 study running 1,280 experimental conditions tested how far content-side properties carry. Read together with C-SEO Bench, the pattern is that evidence helps and the effect is small.
3. Attribution the engine can resolve
57% of AI citations point at sources brands do not own, so a large share of what the answer is built from sits off your domain entirely. Your own name has to resolve cleanly too: entity-oriented retrieval research across 443 configurations shows how much retrieval quality depends on the system knowing which entity a string refers to.
Does structure change what gets extracted, or just what gets read?
It changes what gets extracted, within limits, and the limits are the interesting part. Liu et al. documented that models attend unevenly across a long input context, with material in the middle used least. That finding describes the assembled context. It does not prove that mid-page evidence is penalized, though it does argue against resting a whole page on one buried passage.
Chunk size sets what a retrieval system hands the model. In a 2026 retrieval benchmark built on 1,186 multi-hop chemistry questions, the retrieved units averaged 188 tokens with a 56.26-token standard deviation. That suggests, in that setting, that a tight paragraph is the container a claim has to fit inside, complete with its own attribution. Treat it as a working assumption drawn from one scientific domain. Commercial answer engines have not been measured this way.
Structure is the floor, not the differentiator. Answer-first paragraphs, question-shaped headings and clean HTML make substance extractable, and they produce nothing on a page with no substance to extract. The C-SEO Bench result is exactly what a floor looks like when everyone has already reached it.
What should you stop doing?
Stop stuffing keywords. It scored below the untouched baseline in the study everyone cites for the opposite conclusion, and it degrades the same pages for human readers.
Stop reaching for persuasion framing. Scarcity and exclusivity framing measurably reduces how often a model recommends a product in controlled tests across 10 fictitious products, where the study's results table records recommendation-rate swings as large as 10.60 percentage points from framing changes in either direction. Marketing register is not neutral input to a model.
Stop serving one thing to crawlers and another to people. Text hidden from users but written for machines is cloaking under Google's spam policy, and it puts your organic presence at risk to chase a citation effect nobody has demonstrated.
Stop compressing all of this into one number. The IAB's framework splits AI visibility into Presence, Prominence, Portrayal, and Persuasion because those move independently. A single composite score hides which one you are losing, which is the only thing you needed to know.
How do you test a tactic on your own pages without fooling yourself?
Start by accepting that a single before-and-after comparison proves nothing here. The SparkToro research found AI engines are highly inconsistent when recommending brands, and Search Engine Land's write-up of a separate study reports that AI recommendation lists rarely repeat exactly across runs. A 2026 analysis of brand dynamics in LLM recommendation systems examines the same churn from the model side.
So measure a rate, not an event. Run each buyer question at least ten times across several days and record how often you appear, are cited, and are recommended. A 2026 variance-components study found run-to-run noise large enough to swallow real differences in small samples, and Ronald Sielinski's "Quantifying Uncertainty in AI Visibility" makes the same argument with confidence intervals. The platform medians in that work show how far apart engines sit on the same question.
Then change one thing at a time and keep a holdout. Pick ten comparable pages, add extractable evidence to five, leave five alone, and rerun the same questions on the same cadence for both sets. Without the holdout you cannot separate your edit from a model update or from ordinary drift.
Watch which sources the answers cite, alongside whether you appear at all. Seer found 87% of SearchGPT citations across a 500-citation sample matched Bing's top results at the time of the study. The source list tells you whether the fix belongs on your site or somewhere else.
How does Trovance test what actually gets you cited?
Trovance runs the protocol above as a standing system, on a repeating cycle. You define the market questions your buyers ask, and it runs them repeatedly across AI engines, preserving every answer run as a snapshot with its full context: who was mentioned, who was cited, who was recommended, and which sources carried the answer. The unit of measurement is the appearance rate across many runs.
Preserved snapshots let you compare. When an answer changes, you can see whether the shift came from your page, from a third-party source that started carrying a competitor, or from ordinary run-to-run variance. Answer coverage across a question set shows where you are absent entirely, which is a different problem from being present and losing the recommendation.
Your Brand Core holds the claims you are entitled to make and the proof behind each one, so a recommended action names a specific missing asset. Drafts are produced from approved claims, a person reviews everything before it publishes, and the next analysis cycle reruns the same questions so you can see whether the answer moved. That is the loop the benchmark evidence supports: change something specific, then verify against current answers.
What Trovance will not do is promise a citation, a ranking, or a recommendation. No honest system can. The engines are probabilistic and the sources they trust sit outside your control, and 3 significant conditions out of 54 is what the testing record looks like for tactics sold as certainties. There is no universal visibility score here either, and we are building toward broader engine and source coverage. That work is in progress.
What should you do this week?
Work the three properties in order. Fetch your five most important pages with JavaScript disabled and read what comes back, because that is what the crawler read. Then audit those pages for extractable evidence: specific numbers and named studies, each dated and attributed to a primary source that fifty other pages have not already quoted.
Then check the sources the answers actually cite for your buyer questions, and see how many of them you have any presence in at all. If the answers are built on review sites and community threads that never mention you, no amount of on-page rewriting reaches that surface, and the work belongs in earned coverage instead.
Finally, set the measurement up before you change anything, so this quarter's edits produce evidence. If you want the runs, snapshots and comparisons handled for you, start a free Trovance analysis and see which buyer questions you are absent from, and which sources are carrying the answers instead.
Read the evidence correctly
The most-quoted study in AI search, quoted correctly - what the GEO paper measured and what it did not.
The GEO playbook for getting cited by AI engines - the tactics that hold up once you remove the ones that do not.
Why more evidence can still produce worse AI answers - adding facts is not linear, and the failure mode is specific.
AI agents need evidence, not more content - why publishing volume stopped correlating with being used.
Test it on your own pages
How to track AI citations - building an appearance rate you can actually act on.
How to measure AI search visibility without one score - the dimensions a composite number hides.
Do headings help AI retrieve long documents? - the structural question, answered with retrieval research.
A relevant page can still miss the proof - relevance and evidence are separate failures with separate fixes.
FAQs
What actually gets content cited by AI engines?
Passages carrying a specific, sourced claim, on pages an engine can fetch, from sources it already trusts. The original GEO benchmark reported gains up to 41% over baseline for adding statistics, quotations and cited sources. Rewrites that change tone or keyword density without adding evidence rarely move citation rank in controlled tests.
Do conversational-SEO rewrite tactics work?
Mostly no. C-SEO Bench retested ten published methods across two tasks and six domains and found only 3 of 54 unilateral conditions produced a statistically significant citation-rank gain. Advantages also shrank when every competing document applied the same method, which is the normal condition in any category worth competing in.
Is fact density a real ranking factor for AI citation?
It is a supported correlation, and there is no published threshold you can hit. Adding statistics, quotations and cited sources produced the largest measured gains in the GEO benchmark, and an analysis of 548,534 pages maps which page traits correlate with being pulled into answers. Treat any exact fact-per-word ratio sold as a rule with suspicion.
Does page structure affect what AI engines extract?
Yes, though only as a floor. Models attend unevenly across long input contexts, and retrieved units in one 2026 multi-hop chemistry retrieval benchmark averaged 188 tokens, roughly a tight paragraph. Answer-first paragraphs and clear headings make evidence extractable at that size. They produce nothing on a page with no evidence to extract.
Why does my competitor get cited when my content is better written?
Because writing quality is not the input being measured. Profound found 57% of AI citations point to sources brands do not control, so the answer may be assembled from review sites and roundups that name your competitor. Check the cited sources before rewriting anything on your own domain.
How many runs does it take to know whether a change worked?
At least ten per question, spread across several days, with an unchanged holdout set for comparison. A 2026 variance-components study found run-to-run noise large enough to swallow real differences in small samples. One before-and-after screenshot cannot distinguish your edit from ordinary drift in a probabilistic system.
Can any tool guarantee citations in ChatGPT or Perplexity?
No, and a vendor claiming otherwise is selling against the testing record. Engines are probabilistic and lean heavily on third-party sources. Recommendation lists also rarely repeat exactly across runs. What a tool can honestly do is measure appearance rates over time and show whether a specific change moved them.



