TL;DR
🔎 Ranking buys eligibility, not the slot: in Seer's 500-citation sample, 87% of SearchGPT citations matched Bing's top results, yet the citation is decided after retrieval by other rules.
🧩 The engine rarely runs your query: ChatGPT search rewrites a message into one or more targeted queries, and an AirOps analysis of 548,534 pages found the traits that get a page pulled in describe extractable content, not SERP position.
🎯 Cited is not recommended: the IAB sorts AI-era visibility measurement into 4 groups for brands, and an answer can cite your page while recommending a competitor.
🧪 Rewriting pages into an AI voice is the wrong fix: C-SEO Bench found only 3 of 54 conditions statistically significant, while the GEO study raised source visibility by up to 40% on its own benchmark when statistics and citations were added.
🕸️ The readers are multiplying: AI crawler share of observed crawl traffic rose from 2.2% to 7.7%, and those crawlers read raw HTML while Google renders JavaScript.
🛒 The set forms early: buyers evaluate roughly five vendors and know 3.8 of them before contact.
In Seer's sample of 500 SearchGPT citations, 87% matched Bing's top results, while C-SEO Bench (NeurIPS 2025) found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. Ranking gets you into the pool. It does not decide the slot. A page can hold position three on Google all quarter and never appear once in an AI answer for the same buyer question.
Both systems are behaving correctly. They share a retrieval layer and little else. A ranking is decided per query, against a page of links, for words the buyer typed. A citation is decided per passage, inside a synthesized answer with a few source slots, for a query the buyer never typed at all.
The short answer to the question in the title: ranking helps you get retrieved, and it does not decide whether you get cited. Being findable in conventional search is close to a precondition on the engines that have been measured. What happens after retrieval runs on different rules, and that is where most SEO instinct stops applying. Eligibility is what your SEO buys.
The slot is decided afterward, by other rules, and the stakes belong to the buyer rather than the dashboard. G2 found AI chatbots are now the single largest influence on B2B shortlists, 51% of software buyers now begin research inside an AI chatbot, and one-third of buyers purchased from a vendor they had never heard of before the research started. The assumption that SEO already covers this is why the gap goes unmeasured.
Why does a top-ranking page miss the AI answer?
Because the engine almost never runs the query you rank for. OpenAI documents that ChatGPT search rewrites a user's message into one or more targeted queries before it retrieves anything. A buyer asking for the safest option for a small regulated team produces several attribute-shaped searches instead of one head-term search, so the query your page ranks for is often not the query that runs. Your ranking still applies, but to searches you never targeted.
Then the answer is assembled, not listed. Ten links become one synthesis with a small number of attributable sources, and the API returns those citations as a separate part of the response, not as the answer itself. Position on a results page has no equivalent here. You are either inside the composed answer or outside it.
The model also brings priors. A 2026 analysis of brand dynamics in LLM recommendation systems shows category associations learned in training are sticky and spread unevenly across brands, so retrieval competes with memory. An AirOps analysis of 548,534 pages mapped which page traits correlate with being pulled in, and those traits describe extractable content rather than SERP position.
How does passage-level retrieval change what gets picked?
Ranking grades a page; retrieval grades a chunk of it. The unit an engine lifts is a passage that answers the sub-question, which is why a long, well-optimized page can rank and still lose to a competitor's three-sentence block that states a number outright.
This is where evidence beats structure advice. The GEO study (Aggarwal et al., KDD 2024) found that adding statistics, quotations, and citations raised source visibility by up to 40% on the study's own benchmark, while keyword stuffing did nothing in the same error-bar table. A page written to satisfy a keyword target reads as topically correct and evidentially empty, and the second property is the one being graded.
Entity resolution sits underneath all of it. Entity-oriented retrieval research across 443 configurations shows retrieval quality varies sharply with how entities are resolved. Applied to a brand that appears under three names and describes itself differently on every page, the implication is that retrieval has nothing stable to match against, which is an inference and not a measured result. Rankings tolerate that ambiguity; passage retrieval does not.
Where is SEO still doing the work?
SEO is doing the work at the index. In Seer's sample of 500 SearchGPT citations, 87% matched Bing's top results, so for that engine the index your SEO team maintains is the pool citations were drawn from. No comparable published figure exists for the other engines. Kill your ranking work and you remove yourself from the candidate set before any of this matters.
The mechanics diverge below the index, though. Google renders JavaScript with its Web Rendering Service during indexing, so a client-rendered page can rank. Vercel's crawler research with MERJ documented that GPTBot and its peers read raw HTML instead, which means the same page can rank on Google and arrive at an AI crawler as an empty shell.
Those crawlers are not a rounding error. Cloudflare measured AI crawler share of observed crawl traffic rising from 2.2% to 7.7% while 6sense documented where B2B sites are losing traffic to LLMs. Two crawlers, two rendering assumptions, one page.
What is the difference between a mention, a citation, a recommendation, and a shortlist?
A mention, a citation, a recommendation and a shortlist position are four different outcomes that a ranking report collapses into one row.
A mention is your name appearing in the answer text. A citation is your page used as a source. A recommendation is the engine advising the buyer to choose you. A shortlist position is surviving the narrowing when options get compared.
The distinction is not academic. The IAB's 2026 framework separates Presence, Prominence, Portrayal, and Persuasion for exactly this reason, and its guidance sorts measurement into 4 groups for brands. You can be cited in an answer that recommends someone else, and that outcome looks like a win in any tool counting citations.
The shortlist is where the money is. 6sense found buyers evaluate roughly five vendors and already know 3.8 of them before contact, so the set is largely formed before your funnel sees anyone. Ranking for the category term does not put you in that set. Being the source an answer trusts, on the attribute the buyer asked about, sometimes does.
Which SEO instincts fail on the citation surface?
The rewriting reflex fails first, and the evidence on it is blunt. C-SEO Bench evaluated conversational-SEO methods across two tasks and six domains and found most of them ineffective, with only 3 of 54 unilateral conditions reaching statistical significance. If your plan is to reword existing pages into an AI-friendly voice, the published data says you should expect nothing.
Persuasion tuning may go backwards. Scarcity and exclusivity framing measurably reduced how often an LLM recommended a product in a controlled test across 10 fictitious products. Whether that transfers to a real landing page has not been tested, so treat it as a reason to check your own copy, not as a rule.
And the tactic of serving different content to crawlers is worse than useless. Cloaking is a violation under Google's spam policy, so the shortcut that would let you write one page for engines and another for people costs you the ranking surface you already have. What survives both surfaces is the same thing: specific, checkable claims stated plainly enough to lift.
How do you tell which surface you are actually losing?
Measure them separately, then compare the queries, not the totals. Pull the buyer questions you care about, run each one at least ten times across two or more engines on different days, and record four columns: mentioned, cited, recommended, and which sources the answer used. Set that list beside your ranking report and look at how little the two lists share.
Do not skip the variance step. SparkToro's research found AI engines are highly inconsistent when recommending brands, and Search Engine Land's write-up of the same variance problem reported that AI recommendation lists rarely repeat exactly. A 2026 variance-components study decomposes that noise into its sources, and the practical read is that run-to-run variation can swamp real differences in a small sample. Sielinski's "Quantifying Uncertainty in AI Visibility" makes the same argument with confidence intervals.
Scale matters too. Semrush's AI Visibility Index analyzed 126 million AI search prompts, a query population no keyword tool was built to describe. A small manual sample reads as directional at best.
The protocol also has one structural weakness: it expires the moment you finish it. Models retrain, sources shift, and last quarter's comparison describes a market that has moved.
How does Trovance compare your ranking surface with your citation surface?
Trovance runs that comparison on a repeating schedule instead of once a quarter. You define the market questions your buyers actually ask, and it runs them repeatedly across AI engines as tracked questions, preserving each answer run as a snapshot: who was mentioned, who was cited, who was recommended, and which sources carried the answer. That record makes the divergence visible, because it holds the queries engines actually resolved rather than the keywords you targeted.
From there the diagnosis separates. A brand with strong rankings and no answer coverage usually has a retrieval or extraction problem on its own pages. A brand that appears but loses the recommendation to a competitor whose evidence gets quoted has a proof gap. Because every snapshot keeps its citations, you can see whether the competitor won on their own page or on someone else's.
Your Brand Core holds the claims you are entitled to make and the proof behind each one, so recommended actions name a specific missing asset instead of a keyword: the benchmark that would counter a quoted study, the comparison page a third-party source will never write for you. Drafts are produced from approved claims, and a person reviews and approves everything before it publishes.
What Trovance will not promise is that a citation or a recommendation follows. No honest system can, because the engines are probabilistic, the competitor is publishing too, and the variance documented above is real. There is no universal visibility score here and no control over what a model says. What the analysis cycle does is rerun the same questions after your work ships and show whether the answers moved, so you are steering against current evidence and not a stale snapshot.
What should you do this week?
Start by proving the gap exists in your own data. Take ten buyer questions, run them across two engines ten times each, and put the resulting query list next to your top ranking keywords. If the two lists barely touch, you have been measuring one surface and being discovered on another.
Then sequence the fixes. Start by fetching your key pages with JavaScript disabled to see what an AI crawler reads, then put extractable evidence into the passages that answer buyer questions, since that is the intervention with published support behind it. Keep the ranking work running throughout, because the index it feeds is where retrieval starts.
Be honest about timing. Crawl and rendering fixes can show in answers within weeks, while earning third-party sources takes months. If you want the comparison without the afternoon of manual runs, start a free Trovance analysis and see which surface is costing you the buyer.
Understand the two surfaces
Does GEO replace SEO? - The direct answer to the follow-up question this article raises.
The AI visibility guide for B2B SaaS - The wider framework this comparison sits inside.
Where AI citations come from - Which sources engines actually pull, by industry.
How cross-platform AI agents build product shortlists - What happens after the citation, at the rung that decides deals.
Fix what the gap exposes
AI crawlers don't run your JavaScript - The rendering split between Google and AI crawlers, in detail.
A relevant page can still miss the proof - Why topical fit is not enough for passage retrieval.
How to measure AI search visibility without one score - A measurement setup that survives run-to-run variance.
How to track AI citations - The practical instrumentation for the citation surface.
FAQs
Does ranking on Google get me cited by AI engines?
Partly. Ranking makes you retrievable, and in Seer's sample 87% of SearchGPT citations matched Bing's top results, so search presence is close to a precondition on that engine. It does not decide the citation itself, because engines rewrite the query, retrieve passages instead of pages, and fill a small number of source slots.
Why does a page that ranks well never appear in AI answers?
Usually because the engine never ran your query. OpenAI documents that ChatGPT search rewrites a user message into one or more targeted queries, so retrieval happens against attribute-shaped questions the buyer never typed. The query you rank for is often not the query that runs, and the passage answering it may not exist on your page.
Is SEO work wasted if I care about AI visibility?
No. In Seer's sample, SearchGPT citations came overwhelmingly from the same index your SEO work maintains, and leaving that index removes you from the candidate set. Treat SEO as necessary and insufficient: it buys retrieval eligibility, while extractable evidence and third-party sources decide which candidate an answer actually cites.
Can I rewrite my pages in an AI-friendly style to get cited?
The published evidence says expect little. C-SEO Bench evaluated conversational-SEO methods across two tasks and six domains and found only 3 of 54 unilateral conditions produced statistically significant gains. The GEO study raised source visibility by up to 40% on its benchmark by adding statistics, quotations, and citations, so evidence beats phrasing.
Why do AI crawlers miss pages that Google indexes fine?
Rendering. Google runs JavaScript through its Web Rendering Service during indexing, so client-rendered pages can rank. Vercel's research with MERJ documented that GPTBot and similar crawlers read raw HTML instead, which makes the identical page an empty shell to them. Fetch your pages with JavaScript disabled to check.
What is the difference between a citation and a recommendation?
A citation means an engine used your page as a source. A recommendation means it advised the buyer to choose you. You can be cited inside an answer that recommends a competitor, which counts as a win in most citation trackers. The IAB framework separates Presence, Prominence, Portrayal, and Persuasion for this reason.
How many runs do I need before trusting an AI visibility measurement?
More than one, and ideally ten or more per question across several days. SparkToro found AI engines highly inconsistent when recommending brands, and a 2026 variance-components study decomposes run-to-run noise that can swamp real differences in a small sample. A single answer is a sample, not a measurement.



