TL;DR
🧪 Retesting the standard rewrite methods found only 3 of 54 unilateral conditions produced a significant citation-rank gain, so most on-page tactics are unsupported.
📈 The technique that does have support is evidence: statistics, quotations, and cited sources improved on baseline by 41% in the GEO authors' own benchmark engine, which is a measured ceiling and not a live-engine result.
🕸️ Access beats persuasion: AI crawlers read raw HTML rather than executing JavaScript, and Cloudflare scanned the 200,000 most visited domains for agent readiness.
📚 57% of AI citations point to sources brands do not control, which caps how much any on-page technique can move.
🎲 AI engines are highly inconsistent recommenders, and 4 separate rungs (presence, prominence, portrayal, persuasion) move independently, so one before-and-after check settles nothing.
Most of what circulates as LLM optimization has been retested by people with nothing to sell, and the retesting was unkind. When a research team rebuilt the published conversational-SEO rewrite methods and scored them properly, only 3 of 54 unilateral conditions produced a statistically significant citation-rank gain. That is the number to budget against, because it separates the techniques worth an engineering ticket from the ones worth a shrug. The survivors sort into two groups, and neither group is about phrasing.
The short version: techniques with published support add extractable evidence to a page, meaning statistics, direct quotations, and cited sources. The published KDD version of the GEO study is where that set of methods was defined, and every number attached to it was produced inside a benchmark engine the authors built rather than a live commercial one. Techniques about access are not persuasion at all: a page an engine cannot fetch, a brand it cannot resolve, and a URL that keeps moving lose before any ranking question gets asked. Everything else, mostly rewriting for tone and density, sits in the bucket the retesting could not confirm.
The stakes justify the care. G2 found AI chatbots are now the single largest influence on B2B shortlists, 51% of software buyers now begin research inside an AI chatbot, and 6sense found buyers already know 3.8 of the roughly 5 vendors they evaluate before a seller hears from them. A quarter spent on techniques that move nothing is expensive in a market that forms its list without you.
Which techniques have published evidence behind them?
Two, reliably: adding statistics, and adding quotations attributed to named sources. Both come out of the same benchmark work, and both are about putting evidence a model can lift into the part of the page that answers the question. The GEO preprint reports its best methods improving on the baseline by 41% in source visibility, with the gain concentrated in the evidence-bearing edits rather than the stylistic ones. That 41% is the single measured figure anyone should quote, and it belongs to the authors' own benchmark.
State the conditions or you will misread the result. The evaluation ran on GEO-bench, a query set the authors assembled, scored by a generative engine they constructed over search results. It is not a measurement of ChatGPT, Perplexity, or AI Overviews, and the authors do not claim it is. The KDD proceedings entry and the project's own citation block are the versions to quote from.
Effect size also varies by domain and by method. An open reproduction of the method set records its top single-method gain at +10.60 points on its own scale, which is a real effect and a much smaller one than the headline. Treat the 41% as a ceiling observed under one stated set of conditions, and plan against the smaller reproduction figure.
What does the retesting say about rewriting?
It says most of it does not hold up. C-SEO Bench took the published methods and ran them across two tasks and six domains, a wider test than the original work applied, and the significant results nearly vanished. The methods were scored at a 0.05 significance threshold, so three hits across fifty-four conditions is close to what chance alone would hand you. Those conditions are not independent tests, which makes the comparison illustrative rather than exact.
The second finding matters more for planning. The methods were also scored in the setting where every competitor applies them, and the advantage collapses once adoption is universal, which is the setting a published tactic drifts toward once rivals read the same paper. A technique that only works while rivals ignore it is a timing bet, not a program.
Some rewriting is worse than neutral. Scarcity and exclusivity framing measurably reduces how often a model recommends a product, tested against 10 fictitious products so brand priors could not contaminate the result, and follow-up work extends the same finding. The persuasion language your conversion team likes is a cost on this surface.
Which techniques are about access rather than persuasion?
Three: making the page fetchable as HTML, making the brand resolvable to one entity, and making the answer liftable from a single passage. They pay regardless of which ranking study you believe, because they decide whether your page is a candidate at all. None of them ask a model to be persuaded of anything. They ask it to be able to read you.
Make the page fetchable as HTML
Most AI crawlers do not execute JavaScript. Vercel's crawler research with MERJ documented that they read raw HTML, while Google renders pages through its Web Rendering Service, so a client-rendered page can rank on Google and arrive blank at the engine your buyer asked. Cloudflare scanned the 200,000 most visited domains for agent readiness, and its separate crawler report puts AI crawler share at 2.2% rising to 7.7%. Check your robots rules against the four relevant OpenAI user agents instead of guessing.
Serve crawlers the same HTML you serve people. Diverging by user agent is cloaking under Google's spam policy, and it also makes your own diagnosis unreliable, since you can no longer tell what an engine read.
Make the entity resolvable
A brand that appears under three names and describes itself differently on every page hands a model a string to match instead of an entity to resolve. Entity-oriented retrieval has been benchmarked across 443 configurations. That work is about retrieval systems rather than brand naming, but the mechanism is the same one your brand fails when it resolves to three different strings. One name, one description, one canonical URL per page, and consistent references off-site remove an ambiguity a model otherwise has to settle on its own, which is necessary and not sufficient.
Make the answer easy to lift
Put the answer where retrieval can reach it. Liu et al. showed models use information at the start or end of a long context far better than material buried in the middle, and retrieval operates on passages rather than pages. In one 2026 study over 1,186 multi-hop chemistry questions, retrieved units averaged 188 tokens with a 56.26-token standard deviation; commercial engines do not publish their chunk sizes, so treat that as an order of magnitude rather than a target.
The advice follows from the shape of the pipeline rather than from that measurement. A self-contained paragraph near the top of a section can be lifted whole. A claim that depends on three paragraphs above it cannot. What actually gets cited is covered separately, so this piece will not re-argue it.
Retrievability gates everything else. Seer found 87% of SearchGPT citations across a 500-citation sample matched Bing's top results, and ChatGPT search issues one or more targeted queries before answering, so the query that reaches a search index is what your page has to match. An AirOps analysis of 548,534 pages maps which page traits correlate with being pulled in.
Where does on-page work stop mattering?
At the point where the answer is assembled from sources you do not own. Profound's citation research found 57% of AI citations point to sources outside the brand's control: review sites, comparison articles, community threads. No on-page technique edits a community thread or a review-site category page.
That is not an argument against the access work, which is cheap and mechanical and should happen anyway. It is an argument against treating on-page technique as the whole program when most of the citation surface sits elsewhere. The honest split is to fix access because it is a prerequisite, add evidence because it has support, and spend what remains earning presence in the sources your engines already cite.
How do you know whether a change worked?
You measure a rate across repeated runs. A single recheck the day after you publish reads variance as a result. SparkToro's research found AI engines are highly inconsistent when recommending brands, and Search Engine Land's write-up of the same work reaches the same conclusion. A technique that looks effective in one before-and-after check has not been tested.
Size the sample against the noise. A 2026 variance-components study found run-to-run variation large enough to swamp real differences in small samples, and work on quantifying uncertainty in AI visibility makes the same point with confidence intervals. We treat ten runs of the same question across several days as a working floor; the study prescribes no number, only that small samples are swamped by run-to-run variance.
Measure the right rung too. The IAB's framework separates presence, prominence, portrayal, and persuasion, and those move independently: being mentioned is not being cited, and being cited is not being recommended. The web search tool returns its output in two main parts, one of which carries the citation annotations, which is why citation is the rung you can observe directly.
How does Trovance turn these techniques into decisions?
Trovance turns technique selection into an evidence question. You define the market questions your buyers ask, and it runs them repeatedly across AI engines, preserving every answer run with its full context: which rung your brand reached in each answer, and which sources the answer was built from. That record is what tells you whether your page never entered retrieval at all or entered it and went unquoted.
The record routes you to the right tier of work. Answer coverage that excludes your pages while a competitor's evidence gets quoted points at access and proof, both of which are on-page and both of which you control. Answers assembled entirely from sources you do not own point somewhere else, and no rewriting reaches them. Your Brand Core holds the claims you are entitled to make and the proof behind each one, so a recommended action names the specific statistic or quotation the evidence record says is missing, rather than asking for another page.
Drafts are produced from approved claims and a person reviews everything before it publishes. After an asset goes live, the next analysis cycle reruns the same questions and compares the new answer snapshots against the earlier ones, which is how you learn whether the change moved anything or whether you were reading variance.
What Trovance will not promise is a citation, a ranking, or a recommendation. The engines are probabilistic, the sources shift under you, and the retesting above is the reason nobody should sell a guaranteed outcome from an on-page edit. There is no single visibility score here either; the rungs are reported separately because they fail for different reasons, and we are building toward tighter attribution between an approved asset and the answers that follow it.
What should you do this week?
Sequence by tier. First, fetch your five most important pages with JavaScript disabled and read what comes back; if it is an empty shell, nothing downstream matters. Second, check your robots rules against the named AI user agents and confirm you are not serving crawlers a different page than people get. Third, take the two or three pages that answer real buyer questions and put a specific number, a named source, and a direct quotation into the first paragraph under each heading.
Then stop and measure before doing more. Run ten buyer questions across at least two engines, repeating each question enough times to see its spread; ten repeats is our working floor rather than a published standard. Record mention, citation, and recommendation separately, and read the sources the answers cite. If most of them are third-party, your next quarter belongs to earned coverage rather than another round of rewriting.
What to skip: rewriting tone for authority, padding keyword density, and any tool that scores your prose against an undisclosed model of what engines want. The published retesting does not support it. If you would rather have the measurement running continuously than as a one-off afternoon, start a free Trovance analysis and see which tier your gap is actually in.
The techniques and the evidence
The most-quoted study in AI search, quoted correctly - what the GEO benchmark actually tested.
What gets cited by AI - the passage-level version of the evidence tier.
How to show up in ChatGPT without chasing hacks - the same argument applied to tactics that circulate without support.
The four gates a page clears before it is cited - the sequence technique has to clear first.
Access, entity, and measurement
AI crawlers don't run your JavaScript - the fetchability tier in detail.
A named entity is not a retrieval signal - why entity consistency is necessary and not sufficient.
There is no such thing as an AI visibility score - why the rungs get measured separately.
How to track AI citations - running the repeated measurement that tells you if a change worked.
FAQs
How do I optimize my content so AI engines cite it?
Work in two tiers. Fix access first: serve raw HTML, keep one consistent brand name and one canonical URL, and allow the AI user agents. Then add extractable evidence to the paragraph that answers each heading, meaning a specific statistic, a named source, and a direct quotation. Skip tone rewriting.
Do LLM optimization techniques actually work?
Some do. Evidence-bearing edits have published support, while conversational rewriting mostly does not: a retest of the standard methods across two tasks and six domains found only 3 of 54 unilateral conditions produced a statistically significant gain. Access work pays regardless, because an unreadable page cannot be cited at all.
What did the GEO study actually prove?
That adding statistics, quotations, and cited sources raised source visibility, with a measured best case of 41% inside a benchmark the authors built and an engine they constructed over search results. It was not a measurement of ChatGPT or AI Overviews, and treating that single ceiling as a forecast for a live engine overstates it badly.
Does keyword density or authoritative-sounding phrasing help?
No published result supports either. Retesting the standard methods across two tasks and six domains left only 3 of 54 unilateral conditions significant. Persuasion language can actively hurt: scarcity and exclusivity framing measurably reduces how often a model recommends a product, tested against fictitious products so brand associations could not explain the effect.
How many runs before I can say a technique worked?
We treat ten runs of the same question across several days as a working floor, and more for close calls. No study prescribes that number. A 2026 variance-components study found run-to-run noise large enough to swamp real differences in small samples, so one before-and-after check after publishing tells you almost nothing.
Why do competitors get cited when my page is better written?
Often because the answer was never built from either of your sites. Profound's research found 57% of AI citations point to sources brands do not control, such as review sites and community threads. When that is the pattern, better prose changes nothing and earning presence in those sources is the actual work.
How long before an optimization shows up in AI answers?
Access fixes land as soon as the page is re-fetched, which nobody outside the engine can observe directly. Evidence changes wait on re-crawl and re-retrieval. Earning third-party coverage takes longer still. Nobody can honestly guarantee a citation, and the only way to know is repeated measurement rather than an expected timeline.



