
6 min
Zach Chmael
In This Article
Ask Gemini the same question twice and it cites a different set of sources about 70% of the time. A single number cannot describe that. Here is what can.
Updated
There Is No Such Thing as an AI Visibility Score
Every AI visibility tool sells you a number. Mine says 34. Yours says 61. The number goes up, and someone gets credit; it goes down, and someone explains why.
The number is not describing what you think it is.
This is not a complaint about vendor rigor, and it is not an argument for measuring nothing. It is a claim about what kind of quantity AI visibility actually is. A search ranking is a position in an ordered list, and it holds still long enough to be read. Citation share is a sample drawn from a distribution that changes every time you look at it. Those are different mathematical objects, and treating the second like the first produces confident decisions built on noise.
How unstable are AI citations, exactly?
Unstable enough that two runs of the same query usually cite mostly different sources.
The most rigorous treatment I have found is Ronald Sielinski's "Quantifying Uncertainty in AI Visibility", published to arXiv in March 2026 with a revised version in June. The study ran repeated queries across Perplexity, SearchGPT, and Gemini under two sampling regimes: daily collections over nine days, and high-frequency sampling at ten-minute intervals.
To quantify how much two answers to the same question overlap, the paper used the Jaccard index at the domain level. A value of 1 means both runs cited exactly the same set of domains. A value of 0 means they shared none.
The platform medians came in at 0.29 to 0.31 for Gemini, 0.33 to 0.40 for SearchGPT, and 0.50 for Perplexity. Even Perplexity, the most stable of the three, cited half a different source set on a second run of an identical query.
The identical-citation rate is starker. Two Gemini responses to the same query shared every cited source in 0.01% to 0.10% of cases. For SearchGPT and Perplexity it was 3% to 8%.
This corroborates what SparkToro and Gumshoe.ai found from the brand side, where 600 volunteers running 12 identical prompts nearly 3,000 times got the same brand list twice under 1 in 100 attempts. Two independent studies, different methodologies, same conclusion. A 2026 variance-components analysis put a number on the reliability problem directly, finding brand-ranking reliability near 0.01 for a single answer and concluding that reliability comes from spreading measurement across models and phrasings rather than repeating one prompt.
Why can't you just average it out?
You can, and you should. The problem is what happens to the average when you compare it to someone else's.
Sielinski's central finding is not that the numbers move. It is that bootstrap confidence intervals reveal many apparent differences between domains fall within the noise floor of the measurement process itself.
Put concretely: if your dashboard shows you at 12% citation share and a competitor at 9%, that three-point gap may be entirely inside the error bars. You have not been beaten and you have not won. You have taken two samples from overlapping distributions and read a story into the difference.
The paper also found that citation rankings are unstable across samples throughout the frequently-cited domain set, not just at the margins. Domains that appear often still moved position between runs. So the "we went from 8th to 5th" narrative is usually a narrative.
None of this makes measurement useless. It makes point estimates useless, which is a much narrower claim and a fixable one.
Are the platforms even measuring the same thing?
No, and this is the part that makes blended scores indefensible rather than merely imprecise.
The same study found the platforms return wildly different citation volumes per response: roughly 40 to 43 for Gemini, 20 to 22 for Perplexity, and six to seven for SearchGPT. A citation share of 10% means something structurally different on a platform citing forty sources than on one citing six.
Semrush's analysis of 126 million AI prompts reached a compatible conclusion from a different direction, finding the overlap between brands mentioned and domains cited can be as low as 30% on Gemini. Being mentioned and being cited are separate events, and their relationship varies by platform.
There is a further wrinkle that almost nobody has picked up. Sielinski found SearchGPT exhibits a deterministic layer: for certain domain-query pairings, nine of the frequently-cited domains returned identical citation counts across every sampling job. Some answers are effectively fixed while others remain highly volatile, which makes SearchGPT qualitatively different as a measurement target rather than just noisier or quieter.
A single composite score averages across platforms with different citation volumes, different mention-to-citation relationships, and in one case a mix of deterministic and stochastic behavior within the same engine. The resulting number is arithmetically valid and descriptively empty.
Does this mean tracking is pointless?
No, and the people arguing that are overcorrecting into a worse position.
The strongest counterargument to everything above goes: every measurement discipline uses proprietary methodology, Domain Rating and Domain Authority disagree and remain useful, and what matters is your own trend in one tool over time rather than cross-tool comparison. That argument is largely right, and refusing to measure because measurement is imperfect is how teams end up with no instrumentation at all in a channel that is reshaping their pipeline.
So the honest position is narrower than "scores are fake." It is:
A single run is not a measurement. It is one draw from a distribution.
A point estimate without a sample size or interval is not a metric. It is a reading with unknown error.
Cross-platform blending destroys information that per-platform reporting preserves.
Trends across sufficient repeated sampling are real and worth acting on.
That last point is the one critics of the critics get right. Directional movement measured properly beats no measurement. The failure mode is not tracking. It is tracking once and believing the output. The stakes are not academic: G2's survey of 1,076 B2B buyers found 69% chose a different vendor than originally planned based on chatbot guidance.
What actually stayed stable across all the noise?
Repeat presence. Which is convenient, because it is also the thing that matters commercially.
SparkToro's research found that while position collapsed under scrutiny, visibility percentage held up: some brands appeared in 60% to 90% of responses for a given intent even as their rank bounced around. Repeat presence means something. Exact rank does not.
That gives you a usable hierarchy of what to measure, ordered by how much signal survives the noise:
Presence rate across many runs. How often do you appear at all, for a given intent, on a given platform. This is the most durable signal available.
Citation share with an interval. Reported per platform, with the sample size stated, as a range rather than a point.
Which sources the engine drew on. The citation set tells you where your reputation actually lives, and it is more actionable than your own position within it.
How you are described when you do appear. A mention in a negative or mismatched context is not a win, and a score cannot tell the difference.
Position. Last, and lightly. The research is consistent that ranking within AI answers is close to meaningless.
Notice that four of the five require keeping the raw answers rather than a computed number. That is the practical implication of all this research: the artifact you need is the response itself. The same logic applies on the technical side, where Vercel and MERJ found GPTBot fetched JavaScript in 11.5% of requests and executed none of it. A score cannot tell you a page was unreadable. A fetch can.
What should you ask a vendor?
Four questions, and the answers are usually diagnostic.
How many times do you run each prompt? If the answer is once, the output is anecdotal. Sielinski provides explicit guidance on sample sizes required for interpretable confidence intervals, so this is a solved problem that some tools have chosen not to solve.
Do you report confidence intervals or sample sizes? A number without either is a reading from one thermometer presented as a climate.
Do you blend platforms into one score? If yes, ask how they reconcile a platform citing forty sources with one citing six.
Can I see the raw answers? This is the one I care most about. A tool that discards the response and keeps only the derived number has thrown away the diagnostic material and left you with a figure you cannot interrogate.
I ask that last question first now, because it predicts the others. Tools built to preserve evidence tend to have thought about sampling. Tools built to produce a dashboard number tend not to have.
Where this leaves the number on your dashboard
Use it as a smoke alarm, not a thermometer.
A visibility score can tell you something changed. It cannot tell you by how much, whether the change exceeds the noise floor, which platform drove it, or what to do about it. Those answers live in the distribution and in the raw responses, not in the summary statistic.
This is why Trovance keeps the answers rather than collapsing them into a single figure: the response text, the citations, the platform, the run, and the variance across runs. The score is the least informative thing we could show you, and it is the only thing most tools keep.
The teams that will get this right over the next two years are not the ones with the highest number. They are the ones who can say, with a straight face, how confident they are in it.
FAQs
Is an AI visibility score accurate?
Not as a point estimate. Sielinski's study found citation metrics are samples from a shifting distribution, and that bootstrap confidence intervals place many apparent differences between domains inside the noise floor. A score without a sample size or interval cannot tell you whether a change is real.
How much do AI citations vary between identical queries?
Substantially. Domain-level Jaccard overlap between two runs of the same query had a median of 0.29 to 0.31 on Gemini, 0.33 to 0.40 on SearchGPT, and 0.50 on Perplexity. Gemini returned an identical citation set in only 0.01% to 0.10% of paired responses.
Can I compare AI visibility scores across platforms?
Not meaningfully. The same research found Gemini returned about 40 to 43 citations per response, Perplexity 20 to 22, and SearchGPT six to seven. A given citation share represents a structurally different achievement on each platform, so blending them into one score destroys the information.
Should I stop tracking AI visibility?
No. Refusing to measure because measurement is imperfect leaves you blind in a channel that increasingly shapes buying decisions, and G2 found AI chatbots are now the single largest influence on B2B shortlists. Track properly instead: repeat each prompt many times, report per platform, include intervals, and treat trends across sufficient sampling as the signal.
What is the most reliable AI visibility metric?
Presence rate across many runs. SparkToro found that while position collapsed under repeated testing, visibility percentage held up, with some brands appearing in 60% to 90% of responses for an intent. How often you appear is durable; where you appear in the list is not.
How many times should each prompt be run?
More than once, and enough to produce an interpretable interval. Sielinski's paper gives explicit guidance on sample sizes needed for confidence intervals and recommends repeated sampling across time windows rather than single collections, since distributions shift between runs.
Why do AI answers change if the content hasn't?
Because generative engines are non-deterministic by design. Sielinski separates system-level stochasticity, where the engine itself varies its output and retrieval, from sampling uncertainty in measurement. Both are present, which is why identical queries return different sources minutes apart.

