TL;DR
🧭 51% of software buyers now begin research inside an AI chatbot: every answer you miss is a shortlist forming without you.
🎲 800 million people use ChatGPT, and answers change run to run: measure appearance rate, never one answer.
📊 57% of AI citations point to sources brands do not control, so share of voice is partly decided off your site.
🧪 Only 3 of 54 tested conversational-SEO conditions produced significant citation gains, so phrasing tricks do not buy share.
📐 The IAB splits AI visibility measurement into four groups rather than one score; a universal number hides where you lose.
Run the same buyer question through ChatGPT a dozen times this week and you will get vendor lists that disagree with each other. SparkToro's research found AI engines are highly inconsistent when recommending brands, and a separate study found AI recommendation lists rarely repeat exactly. That instability is the first fact of this metric: any share-of-voice number built from one answer per question is noise wearing a percentage sign. The honest metric is an appearance rate, measured across repeated runs.
The definition fits in two lines. Appearance rate: answer runs naming your brand, divided by total runs of that question. AI share of voice: your total appearances divided by all tracked brands' appearances in the same runs, per engine, over a stated window. The stakes justify the discipline: G2's research found AI chatbots are now the single largest influence on B2B shortlists, 51% of software buyers now begin research inside an AI chatbot, and buyers evaluate roughly five vendors with most of the list set before they ever contact sales.
Why is one answer per question not a measurement?
Because the engine is probabilistic, a single answer is a sample, and appearance rates are the only stable ground. A 2026 variance-components study found run-to-run noise large enough to swamp real brand differences in small samples, and work on quantifying uncertainty in AI visibility shows how wide the confidence interval around a small-sample rate is. Ten runs per question is the floor.
This is where AI share of voice parts company with the search version. Traditional share of voice is a position in a ranked list many brands occupy at once; an AI answer names a short list of vendors and everyone else is absent. Presence is close to binary, which makes the appearance rate, and who appears alongside you, the unit of competition.
Keep the rungs separate from the start: a mention is your name in the answer, a citation is your page used as a source, and a recommendation is the engine advising the buyer to pick you. Profound's citation research found 57% of AI citations point to sources brands do not control, so your citation count and mention count can tell opposite stories.
How do you calculate AI share of voice?
Count appearances across repeated runs, then divide twice. Per question, appearance rate equals runs naming your brand over total runs. Across the panel, share of voice equals your total appearances over all tracked brands' appearances in the same runs. Log citations and recommendations as separate columns; a weighted composite has arbitrary weights and destroys the diagnosis.
A worked example: five buyer questions, each run twelve times on one engine in a month, sixty answer runs. You appear in 21 runs, Competitor A in 48, Competitor B in 33, Competitor C in 12: 114 tracked appearances in total.
| Brand | Appearances in 60 runs | Appearance rate | Share of voice |
|---|---|---|---|
| Competitor A | 48 | 80% | 42% |
| Competitor B | 33 | 55% | 29% |
| You | 21 | 35% | 18% |
| Competitor C | 12 | 20% | 11% |
The single number says 18% and third place. The per-question split says more: those 21 appearances were 11 of 12 on one question, 8 of 12 on another, and 2 of 36 across the remaining three. A fixture on one question, solid on a second, near absent on the other three: that split is a work plan, while 18% is only a scoreboard.
Two rules keep the number honest: state the window and run count with every report, and never compare your number against one built from a different panel, engine mix, or counting rule, including a vendor's.
Which questions and competitors belong in your measurement panel?
The questions a real buyer would ask before choosing, and the competitors the engines actually surface: both are empirical, and neither comes from your positioning deck. Semrush's AI Visibility Index analyzed 126 million AI search prompts, a reminder that the space of real prompts is enormous and any panel is a deliberate sample. Write five to ten questions in buyer language: category, comparison, and use-case questions, skipping the brand-name lookups you already win.
Build the competitor set from the answers rather than from instinct. Run the panel once, log every vendor the engines name, and include brands that keep appearing alongside you even if sales never meets them. Buyers meet vendors this way: one-third of buyers purchased from a vendor they had never heard of before an AI introduced them, and a set that omits those adjacent brands inflates every share number in the report.
Size the set for stability. Two or three brands make the metric jumpy, because one competitor launch swings every ratio; fifteen dilute it with vendors no buyer weighs. Commercial pressure on these answers is rising: commercial conversations in ChatGPT more than doubled in a year in a 7.5 million-conversation sample, and Gartner predicted traditional search volume would drop 25% by 2026 as chatbots absorb queries.
What does a single universal score hide?
It hides the splits that tell you where you are losing, the entire point of measuring. Engines retrieve differently: Seer found 87% of SearchGPT citations matched Bing's top results, so a brand strong in Bing's index can lead on one engine and trail on another, and the same uncertainty research reports platform medians that differ by engine. Average across engines and the platform where you are losing disappears into the platform where you are winning.
The same erasure happens across rungs. A score that blends mentions, citations, and recommendations with fixed weights asserts that a citation is worth some exact multiple of a mention, a number nobody has measured. The IAB's guidance on measuring visibility in the AI era splits brand measurement into four groups, Presence, Prominence, Portrayal, and Persuasion, because one number cannot carry them all.
This is the vendor-theater test. A dashboard selling one universal score is selling comfort: a blended number cannot fall on the question that funds your quarter while rising overall. The honest report is small and split: appearance rate per question, per engine, per rung, with run counts attached. If a tool cannot show you the runs behind its number, treat the number as marketing.
How do you read movement without fooling yourself?
Test any change against run-to-run noise first, because an appearance rate that moves within its variance band has not moved. Small panels make this worse: at ten runs per question, one flipped answer moves a rate by ten points. Before celebrating or escalating, re-run the panel and see whether the movement survives.
Then read absolute and relative together. Your appearance rate can climb while your share of voice stays flat, which means the category is getting more visible and you are only keeping pace. Only the pair tells you whether your work moved you or the market moved everyone.
Falling share despite good work has three usual causes: a competitor shipped the evidence the engines now quote, a new entrant expanded the set, or an engine changed retrieval. A 2026 analysis of brand dynamics in LLM recommendation systems found prior brand associations sticky and unevenly distributed: slow gains are normal, and a sudden jump in either direction deserves suspicion before it earns a slide in the board deck.
What actually moves AI share of voice?
Evidence that engines can retrieve and quote, on your pages and in the sources they cite; phrasing tricks are the part that does not work. The GEO study (Aggarwal et al., KDD 2024) found adding statistics, quotations, and citations lifted citation visibility by roughly 30 to 40% in its benchmark. The same literature carries the corrective: C-SEO Bench (NeurIPS 2025) tested ten conversational-SEO rewrite methods and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. Add evidence buyers can check; do not rewrite prose for machines.
Some language actively hurts: scarcity and exclusivity framing measurably reduces how often an LLM recommends a product. Comparison and alternatives questions are the highest-value targets: they force the engine to name several vendors, so your absence is visible in every run. A comparison page that states checkable claims about you and your rivals gives the engine something quotable when those questions arrive.
And spend where the citations are. With a majority of AI citations pointing at third-party sources, review sites and community threads carry more of the answer than your homepage does, so appearance gaps are often closed off-site. Match the work to the rung you are losing: retrievability if you are absent, on-page evidence if you are mentioned without being cited, earned sources if a rival owns the answers.
How does Trovance measure your share of voice against competitors?
Trovance runs the protocol in this article as a standing system instead of a monthly spreadsheet exercise. You define tracked questions in buyer language, and it runs them repeatedly across AI engines, preserving every answer run as a snapshot with its full context: which brands were mentioned, which were cited, which were recommended, and which sources carried the answer. Appearance rates and share of voice are computed from those preserved runs, per question and per engine, with run counts visible.
Because the raw runs are kept, the competitive picture stays inspectable. Answer coverage shows where each competitor appears and where you are absent, and any number can be traced back to the snapshots behind it. When a rival holds a question, the citations in those runs show whether they won it on their own pages or on a third-party source, which is the difference between a content task and an earned-media task.
Measurement then feeds work. Your Brand Core holds the claims you are entitled to make and the proof behind each one, and recommended actions name the specific asset the evidence record says is missing for the question you are losing. Drafts are produced from your approved claims, a person reviews everything before it publishes, and the next analysis cycle reruns the same questions so you can verify whether the asset moved the rate.
What Trovance will not do is report a universal score or promise a share-of-voice target. The engines are probabilistic, competitors keep publishing, and a single number blended across engines and rungs would hide exactly the failures this article describes. What it offers instead is a preserved, comparable record: the same questions, run after run, cycle after cycle, so the movement you act on actually happened.
What should you do this week?
Build the smallest honest panel and run it. Write five questions a real buyer would ask before choosing in your category. Run each ten times across ChatGPT and one other engine, on at least three days, logging mentions, citations, and recommendations separately. Compute appearance rates per question and share of voice across the set you actually observed.
Then resist the summary. Do not average engines or blend rungs into a composite, and put the run count on every number you circulate. Where the split shows a question you never appear on, that is next month's work; where it shows variance, the correct response is another measurement cycle, not a strategy meeting.
Be honest about pace. Appearance rates move over weeks and months as engines re-crawl and re-retrieve, and anyone promising a share-of-voice target on a schedule is selling against the evidence above. If you would rather define the questions once and let the runs accumulate, start a free Trovance analysis and measure your share of voice from preserved answer runs instead of a one-off spreadsheet.
Measure without the theater
How to measure AI search visibility without one score: the framework this metric belongs inside
There is no such thing as an AI visibility score: why a universal number fails on contact with data
The prompt panel is measuring the wrong question: how to choose buyer questions worth tracking
The 10-minute audit of what ChatGPT tells buyers about you: the fastest first look before you build a panel
Close the gaps you find
Why ChatGPT recommends your competitor: the five failure modes behind a losing share number
Where AI citations come from across industries: the third-party sources that decide answers off your site
How to build a competitor comparison page that survives scrutiny: the asset comparison questions keep asking for
A visibility gap is not a content brief: deciding what the measurement tells you to do
FAQs
What is AI share of voice?
AI share of voice is the percentage of repeated AI answer runs in which your brand appears, measured for a defined set of buyer questions, engines, and competitors over a stated time window. It is an appearance rate across many runs, never a judgment from a single answer, because single answers do not replicate.
How many times should I run each question to measure it?
Ten or more runs per question, spread across several days, is a defensible floor. Research on AI visibility measurement found run-to-run variance large enough to swamp real differences in small samples, so fewer runs produce appearance rates that move for no reason. More runs narrow the uncertainty around every rate.
Should mentions and citations be combined into one score?
No. A mention is your name in an answer; a citation is your page used as a source; a recommendation is the engine advising buyers to choose you. Each moves for different reasons and responds to different work, so blending them into one weighted score hides which rung you are losing.
Which AI engines should I measure share of voice on?
Measure each engine your buyers use separately, starting with ChatGPT plus at least one other. Engines retrieve differently: 87% of SearchGPT citations matched Bing's top results in one study, and published platform medians for visibility metrics differ by engine. An average across engines hides the platform where you are losing.
What is a good AI share of voice number?
There is no universal benchmark. AI share of voice only has meaning relative to the competitor set and question panel it was measured against, and every one of those choices changes the number. Track your own trend on a fixed panel over time and compare rivals within one measurement, never across vendors' dashboards.
How often should I re-measure AI share of voice?
Monthly is a practical cadence for most teams, with weekly runs reserved for active campaigns or launches. Models retrain and competitors publish, so any snapshot expires. Re-run the same question panel on the same engines each cycle so that movement reflects the market rather than a changed measurement.
Can prompt-optimization tactics increase my share of voice?
Rarely. C-SEO Bench tested ten conversational-SEO rewrite methods and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains, and scarcity-style persuasion language measurably reduced recommendation rates in separate testing. Appearance gains come from retrievable pages, extractable evidence, and presence in the sources engines cite.



