TL;DR
🛒 The shortlist forms before you hear about it: 51% of software buyers now begin research inside an AI chatbot.
🧭 Buyers evaluate 3.8 of the roughly 5 vendors they consider before they contact anyone, so being absent from the answer costs you the meeting itself.
🎲 Nobody can guarantee an outcome: a 2026 variance-components study found run-to-run noise large enough to swamp the brand differences being measured, so every honest reading is a rate across many runs.
📚 57% of AI citations point to sources brands do not control, which puts part of this in the communications budget.
🧪 C-SEO Bench found only 3 of 54 tested conditions produced significant citation gains: buy a baseline before you buy tactics.
By the time a buyer talks to your sales team, the shortlist usually exists already, and more of it is now assembled inside an AI assistant that leaves no trace in your analytics. G2 found AI chatbots are now the single largest influence on B2B shortlists, 51% of software buyers now begin research inside an AI chatbot, and one-third of buyers purchased from a vendor they had never heard of before they started researching. That is the case for spending ten minutes on this.
Here is the argument, compressed. Buyers form shortlists inside assistants. Being named in an answer is a different asset from ranking on a results page, and it fails for different reasons.
The measurement is probabilistic, so any vendor promising a guaranteed citation or ranking is selling against the published evidence. Most of what an engine cites sits on properties you do not own, which puts a share of this work in the communications budget. And the correct first move is a baseline you can repeat.
What actually changed in how your buyers choose?
The research phase moved into a conversation, and it now happens earlier with fewer of your pages involved. 6sense found buyers evaluate roughly five vendors and have most of that list in mind before contacting anyone, and 58% of buyers say they engaged sellers sooner. The plausible reading is that the early comparison work was already done for them.
Volume is not the interesting number, but it sets the scale. In April 2025, Sam Altman put ChatGPT's reach at roughly 800 million people, and Profound's 7.5 million-conversation sample found commercial conversations in ChatGPT more than doubled in a year. Commercial intent inside a conversation falls well short of buyer demand, and the direction still tells you the early evaluation is happening somewhere you cannot see.
What this does not mean is that your website stopped mattering. It means the first read of your company often happens somewhere you cannot instrument, and the buyer arrives holding an opinion that someone else assembled.
Why is being in the answer a different asset from ranking?
Because the outcomes people lump together are separate rungs with separate failure modes. A mention is your name appearing in an answer. A citation is your page used as a source. A recommendation is the assistant advising the buyer to pick you.
A shortlist position is surviving the narrowing when the buyer asks it to compare. You can hold the lower rungs and lose the one that produces revenue.
Retrieval and memory both feed the answer. Seer found 87% of SearchGPT citations across a 500-citation sample matched Bing's top results, so ordinary search visibility still gates much of what gets read. But a 2026 analysis of brand dynamics in LLM recommendation systems shows the associations a model already carries are sticky and spread unevenly across brands. Retrievability is necessary and not sufficient.
If you want one vocabulary to hold your team to, the IAB's 2026 guidance on measuring visibility in the AI era splits the problem into presence, prominence, portrayal and persuasion. The value of that split is budgetary. Each one breaks for a different reason and gets fixed by a different team.
Why should you distrust any guaranteed outcome?
Because the same question asked twice can produce different vendors, so every honest measure of AI visibility is a rate with a sample size attached. SparkToro's research found AI engines are highly inconsistent when recommending brands, and a separate study found AI recommendation lists rarely repeat exactly.
The academic work says the same thing with error bars. A 2026 variance-components study found run-to-run noise large enough to swamp real differences in small samples, and work on quantifying uncertainty in AI visibility puts confidence intervals around the estimates. One screenshot of a bad answer is a single sample.
Two commercial consequences follow. A vendor promising guaranteed citations, guaranteed recommendations or control over what a model says is promising something the published research says is not available. And a single blended visibility score hides the rung you are actually losing, which is the number your team will end up managing to. Ask instead for appearance rates on named buyer questions, with the sample size attached.
Why is this only partly a content problem?
Because most of what the engine repeats back does not live on your domain. Profound's citation research found 57% of AI citations point to sources brands do not control: review platforms, comparison articles, community threads, trade coverage. When those sources name a competitor and omit you, the model is reporting its sources faithfully.
That reframes the budget question. Some of the work belongs to communications, analyst relations and customer proof, and only some belongs to your content calendar. For the smaller share you do own, an AirOps analysis of 548,534 pages maps which page traits correlate with being pulled into an answer.
The mechanical half deserves one line of attention because it is cheap to check. Vercel's crawler research with MERJ documented that the OpenAI and Anthropic crawlers do not execute JavaScript, while Google's crawlers bring the same Web Rendering Service capabilities to the job. Cloudflare tested the 200,000 most visited domains for agent readiness. Ask your engineering lead whether your key pages render in raw HTML: five-minute question, expensive answer.
What does the evidence on tactics actually support?
Less than the vendor decks claim, and the strongest published result is a corrective. C-SEO Bench, published at NeurIPS 2025, found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. The benchmark tested conversational-SEO methods across two tasks and six domains. Most of what gets sold as rewriting for AI search does not survive that test.
What holds up is evidence on the page. The GEO study's strongest method improved on its baseline by 41% inside that paper's own benchmark, which is a result about one measured method and not a promise about your category. Brief your team accordingly: defensible numbers and named sources on the pages that answer buyer questions, and no phrasing tricks.
Some persuasion language is worse than neutral. Scarcity and exclusivity framing measurably reduces how often a model recommends a product. The voice that converts on a landing page can cost you inside an answer, which is a copy decision your CMO should be making deliberately.
What is the correct first move?
A baseline comes first. Write down the five to ten questions a real buyer asks before choosing in your category, run each of them repeatedly across at least two assistants over several days, and record who is mentioned, who is cited, who is recommended, and which sources carried the answer.
That panel is a proxy. Research across 670 English commercial multi-turn conversations shows buyers reach a shortlist over an exchange, so read your panel as a rate you track over time and expect real buyers to arrive by a longer road.
The exercise still narrows the questions worth settling this quarter: whether you have a real gap or run-to-run variance, whether the gap sits on your pages or on third-party sources, and whether the engine resolves you as a distinct company at all. Entity-oriented retrieval research across 443 configurations shows how much retrieval quality depends on that last point, and a brand that appears under three names never resolves cleanly.
Market-scale readings exist if you want context. Semrush's AI Visibility Index analyzed 126 million AI search prompts. The version that changes your decisions is built on your buyers' actual questions, which is a shorter and far more specific list than any index.
How does Trovance give a CEO a defensible read?
Trovance runs that baseline continuously. You define the market questions your buyers ask, and it runs them repeatedly across AI engines, preserving each answer run with its full context: who was mentioned, who was cited, who was recommended, and which sources the answer was built on. The record is the product, because one answer is thin evidence and a quarterly screenshot expires.
Because every answer snapshot keeps its citations, the diagnosis separates itself. Answer coverage that never includes you while your pages resist a raw fetch is a retrieval problem for engineering. Answers assembled entirely from third-party sources that omit you are a communications problem. Answers that describe a different company are an entity problem.
From there the work is scoped by evidence. Each diagnosis routes to a different owner and a different budget line, which is the distinction a CEO is actually buying. Your Brand Core holds the claims you are entitled to make and the proof behind each one, so a recommended action names the specific missing asset: the benchmark that would counter a competitor's quoted study, the comparison page a displaced third-party source will never write for you. Drafts are produced from approved claims, and a person reviews everything before it publishes.
What Trovance will not promise is a guaranteed citation, a guaranteed recommendation, or a single number that says how visible you are. The engines are probabilistic, your competitors are publishing too, and a blended score hides the rung you are losing. What it does instead is close the loop. After an asset ships, the next analysis cycle reruns the same questions and shows whether the answers moved, so you are deciding against current evidence.
What should you do this month?
Three asks, in order. Ask your marketing lead for appearance rates on ten named buyer questions with the number of runs behind each rate. Ask engineering whether your key pages render without JavaScript. Ask whoever owns communications which third-party sources the answers cite, and whether you have any presence in them.
Then set the expectation honestly with your board. Our working expectation, which you should audit instead of taking on faith, is an ordering: retrieval fixes surface fastest, evidence added to your own pages follows them, and earned third-party coverage moves slowest of the three. Every reading you take carries the variance documented above. Fund a program that measures, not a campaign that promises.
If you want the baseline without an afternoon of manual runs, start a free Trovance analysis and see which rung your category is actually losing.
Get a baseline first
How to see what ChatGPT says about your company - the fastest first reading you can take yourself.
How to measure AI search visibility without one score - what to ask your team for instead of a dashboard number.
The prompt panel is measuring the wrong question - the limits of a repeated single-turn baseline.
There is no such thing as an AI visibility score - why a composite number hides the rung you are losing.
Decide where the budget goes
Where do AI citations come from - how much of the answer sits off your own domain.
How to earn third-party AI citations - the communications half of the work.
Does AI search visibility drive leads or revenue - the question your CFO will ask second.
A visibility gap is not a content brief - why a gap does not automatically justify a page.
FAQs
What does a CEO need to know about GEO and AI search?
Five things. Buyers now build shortlists inside AI assistants before they contact vendors, being named in an answer is a separate asset from ranking, and the measurement is probabilistic, so appearance is a rate across runs. Most cited sources sit on properties you do not own, and the first move is a repeatable baseline on your buyers' real questions.
Is GEO a replacement for SEO?
No, and treating it as one wastes budget. Seer found 87% of SearchGPT citations across a 500-citation sample matched Bing's top results, so search visibility still gates much of what an assistant reads. What changes is the outcome you measure: appearance and recommendation rates inside answers, always reported with sample sizes.
Can any vendor guarantee my brand appears in ChatGPT answers?
No. SparkToro found AI engines are highly inconsistent brand recommenders, and a 2026 variance-components study found run-to-run noise large enough to swamp real differences in small samples. Guaranteed citations, guaranteed recommendations and control over model output are all promises the published evidence does not support. Treat them as a disqualifier.
How much of AI visibility is really a content problem?
Less than half of it by citation volume. Profound found 57% of AI citations point to sources brands do not control, including review platforms, comparison articles and community threads. That share gets earned through communications, analyst relations and customer proof, which sits in a different budget line from your content calendar.
Do the popular GEO tactics actually work?
Mostly not. C-SEO Bench found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains, across two tasks and six domains. The GEO study's strongest method improved on its baseline by 41% inside that paper's benchmark. Brief your team on defensible numbers and named sources instead of buying rewrite tactics.
What should I ask my marketing team first?
Ask for appearance rates on ten named buyer questions, with the number of runs behind each rate. Then ask which third-party sources the answers cite and whether you appear in any of them. A dashboard score without a sample size is not an answer to either question.
How long before AI answers about my company change?
It depends on the failure mode, and our working assumption is an ordering with no promised date. A retrieval fix moves fastest, evidence added to your own pages follows it, and earned third-party coverage is slowest of the three. Nobody can honestly promise a date, because measurement variance alone makes short-window claims unfalsifiable.



