TL;DR
🗺️ A citation audit produces three regions: cited, rival-cited, and unowned. 57% of AI citations point to sources brands do not control, so most of the map is not a content brief.
🎯 Size the gap before you queue it: a 2026 variance-components study found run-to-run variance in AI visibility measurement, which is why a single run is not a reading.
🔎 Retrievability gates everything. 87% of SearchGPT citations matched Bing's top results, which makes an engine-specific gap a retrieval question before a content one.
🧪 Rewrites rarely move the rung you care about: C-SEO Bench found only 3 of 54 tested conditions produced statistically significant citation-rank gains.
📊 Evidence does move it. The GEO study at KDD 2024 found statistics, quotations, and citations raised citation visibility by roughly 30 to 40% on its benchmark, while keyword stuffing did nothing.
Run twenty buyer questions ten times each across two AI engines and you get a map with three regions: questions where your brand is cited, questions where a rival is cited instead, and questions where nobody in your category is cited. The second region is the one teams act on first, and that is backwards, because a gap turns into a content brief only when the citation that would fill it could land on a page you own.
Most of them cannot. Profound's citation research found 57% of AI citations point to sources the brand does not control: review sites and community threads. The audit still earns its afternoon: G2 found AI chatbots are now the single largest influence on B2B shortlists, 51% of software buyers now begin research inside an AI chatbot, and one-third of buyers purchased from a vendor they had never heard of before. The waste happens in the next step, when a spreadsheet of absences becomes a publishing calendar.
What does a citation gap map actually show?
Record three fields per question: who was named, whose page was used as a source, and which domain that source sits on. Collapse them into a single brand figure and you get a number that moves without telling you where to work. The third field decides whether any of this is your job.
The three regions behave differently enough that treating them alike is the core mistake. Where a rival is cited, start with where the citation lives, because a competitor quoted from their own documentation and a competitor named inside a review roundup are unrelated problems. Where nobody in the category is cited, the engine assembled the answer from generic material, and that is the only region with open ground in front of it.
Keep the rungs separate while you record. A mention is your name in prose. A citation is your page carrying the answer, and a recommendation is the engine telling the buyer to choose you. The IAB's 2026 measurement guidance separates presence, prominence, portrayal, and persuasion for the same reason, and groups those measures into four sets for brands. Merging them back into one figure is a choice the guidance leaves to you, and it is the wrong one: a gap on one rung does not imply a gap on the others.
How many runs before a gap counts as a gap?
Enough that the difference you are about to spend money on exceeds the noise between runs. SparkToro's research found AI engines are highly inconsistent when recommending brands, with the same prompt returning different vendor lists on repeat, a result Search Engine Land summarized as recommendation lists rarely repeating exactly. A single absence is a sample, not a finding.
Treat every cell as a rate with an interval around it. A 2026 variance-components study found run-to-run variance in AI visibility measurement, and work on quantifying uncertainty in AI visibility reaches the same conclusion by putting confidence intervals on visibility measures. Our own floor is ten runs per question across several days.
The rule that falls out of this: rank the queue by the size of a gap relative to its own variance. Appearing in one run of ten while a rival appears in eight is worth a decision. Trading places between Tuesday and Thursday is worth another month of observation and nothing else. Category-level claims rest on corpora like Semrush's AI Visibility Index and its 126 million AI search prompts. A twenty-question panel is a different instrument, and it should not be read as one.
Record the engine and the date with every run. Answers move as models retrain and as the sources behind them change, so an undated map cannot tell you whether a gap closed.
Which questions belong in the set you measure?
The questions a buyer asks before choosing, in the words they use. The keywords you already rank for are a different list, built for a different machine. In our experience ten to thirty questions is where a weekly cadence survives contact with real work; past that, the running eats the analysis. Mix category questions, direct comparisons, and job-to-be-done phrasings that never name a vendor.
The three types fail differently, which is why the map is useful. Comparison questions tend to be answered from third-party sources. Job-to-be-done questions go unanswered by name most often, which is where an unowned region opens. Category questions are where prior associations dominate, and a 2026 analysis of brand dynamics in LLM recommendation systems shows those associations are sticky and spread unevenly across brands.
Size the set against how buyers actually decide. 6sense found buyers engage 3.8 of the roughly five vendors they consider, so the questions that matter are the ones that assemble that list. Profound's 7.5 million-conversation sample found commercial conversations in ChatGPT more than doubled in a year, so their supply is growing faster than most panels are.
How do you tell an addressable gap from one that is not?
Read the citations under the answer, then ask whether a page you could publish would have been eligible to appear there. That single test does more triage than any scoring model, and it is the reason a visibility gap is not by itself a content brief. The absence is evidence that something is wrong; it is not evidence that the fix is an article.
Four checks, in order, resolve most of the map. First, is the answer built on domains you cannot publish on? If a review platform and two forum threads carry it, your site was never in the running, and Profound's finding that 57% of citations sit on sources brands do not control describes that shape.
Second, can the engine read your pages? Vercel's crawler research with MERJ documented that the AI crawlers it measured did not execute JavaScript, and Cloudflare's agent-readiness analysis of the 200,000 most visited domains found much of the web unreadable to agents. That is an engineering ticket, not a brief.
Third, does the engine resolve you as a distinct company? Entity-oriented retrieval research across 443 configurations shows how much retrieval quality depends on that resolution, and a brand split across three names will not be cited consistently for anything. Fourth, and only then: is there a claim you can prove that the current sources do not carry? That is the addressable gap, and it is a minority of the map.
Be skeptical of the tactic layer while you do this. C-SEO Bench (NeurIPS 2025) tested conversational-SEO methods across two tasks and six domains and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. What does move is evidence: the GEO study published at KDD 2024 found that adding statistics, quotations, and citations raised citation visibility by roughly 30 to 40% on its benchmark, while keyword stuffing did nothing. Some framing actively hurts, since scarcity and exclusivity framing measurably reduced recommendation rates in a controlled test on fictitious products.
How should the queue be ranked?
By addressability first, then by evidence you already hold, then by traffic value. Reversing that order is how teams publish four articles against a gap a review site owns. A row should fall the moment you learn its answer was assembled somewhere you cannot publish.
Three inputs decide the top of the list. Whether the question is retrievable ground for you, which Seer's finding that 87% of SearchGPT citations matched Bing's top results makes concrete: if nothing of yours surfaces, the citation is decided among pages that did. Whether you hold provable material, since AirOps mapped the page traits correlated with being cited across 548,534 pages and those traits are structural. And whether the gap survived the variance test.
Everything else belongs in one of two side lanes. Third-party displacement goes to whoever owns reviews and analyst relations, with the domains named from your own citation record. Fit gaps, where the rival is honestly the better answer, go to product and positioning. Neither is a failure, and recording both as results is what keeps the queue short enough to finish.
How does Trovance turn a citation gap into a queue you can defend?
Trovance runs your tracked questions repeatedly across AI engines and preserves each answer run as a snapshot with its citation record, so the map persists as a standing artifact across cycles. Answer coverage across that record produces the three regions. Because every snapshot keeps its sources, you can see whether a competitor won on their own documentation or on a domain neither of you controls.
Variance is handled by repetition. Runs accumulate across cycles, so a gap is compared against its own history before it counts as a finding, and a question that flips week to week reads as instability. Your Brand Core holds the claims you are entitled to make and the proof behind each one, which is what turns an addressable gap into a specific brief.
Recommended actions carry the reason they exist, including the reason not to write. When the record shows an answer was assembled from third-party domains, the action names those domains instead of proposing a page, which keeps the queue honest about what publishing can reach. Drafts are produced from approved claims, and a person reviews and approves everything before it ships.
What Trovance will not promise is that a queued brief earns the citation. No system can, because engines are probabilistic and the sources they trust belong to other people. There is no universal visibility score here and no control over what a model says. The analysis cycle reruns the same questions after work ships, so the next decision is made against current evidence.
What should you do this week?
Start with twenty buyer questions and ten runs each, split across two engines and three days. Record mention, citation, source domain, and the rival named, and resist acting until the sheet is full. The first pass is a base rate, and everything downstream depends on it being one.
Then sort. Questions whose answers rest on domains you cannot publish on go to a third-party lane, and questions where your pages exist but were unreadable become engineering tickets. Questions where the rival is the honest answer go to positioning. What survives is the queue, and it will be shorter than the audit implied.
Be plain about timing. Expect retrieval fixes to resolve faster than earned third-party presence, and expect the measurement noise above to stay. Neither interval is something anyone can quote you in advance. If you would rather not run the panel by hand each week, start a free Trovance analysis and let the citation record build itself while you work the queue.
Measure the gap before you queue it
A visibility gap is not a content brief - the argument this method builds on.
How to track AI citations - the logging discipline behind a trustworthy map.
The prompt panel is measuring the wrong question - why question selection decides what you can see.
How to measure AI search visibility without one score - what to report when one composite hides the decision.
Decide what the gap justifies
Where AI citations come from - how much of the answer sits on domains you never own.
What gets cited by AI - the page traits that make a brief worth writing.
AI agents need evidence, not more content - why proof beats volume on an addressable gap.
How to build a competitor comparison page - the asset a rival-cited gap most often justifies.
FAQs
What topics do AI engines cite my brand for, and not?
Only a repeated run of your buyer questions can answer that. Log every answer three ways: who was mentioned, whose page was cited, and which domain carried it. Ten runs per question across several days produces rates you can trust, because the same prompt returns different vendor lists on repeat.
How large does a citation gap need to be before I act on it?
Larger than the noise between runs on that same question. A 2026 variance-components study documented run-to-run variance in AI visibility measurement, and our own floor is ten runs per question. If you appear in one run of ten and a rival appears in eight, that is a finding. If you swap places weekly, keep watching.
Does a citation gap mean I should write a new page?
Usually not. Profound's research found 57% of AI citations point to sources brands do not control, so many gaps are held by review platforms, roundups, and forums where publishing on your own site changes nothing. Write only when the missing evidence could plausibly live on a page you own.
How many questions should a citation gap audit track?
In our experience, ten to thirty is the sustainable range for a weekly cadence, mixing category questions, head-to-head comparisons, and job-to-be-done phrasings that name no vendor. Beyond that, running the panel consumes the time meant for analysis. Cap the set, keep it stable for a quarter, and expand it only when a decision needs it.
Can rewriting existing pages close a citation gap?
Rarely on phrasing alone. C-SEO Bench found only 3 of 54 tested conversational-SEO conditions produced statistically significant citation-rank gains. The GEO study at KDD 2024 found something different works: adding statistics, quotations, and citations lifted citation visibility roughly 30 to 40% on its benchmark. Evidence moves, wording mostly does not.
Why does my brand get cited on some engines but not others?
Because retrieval differs before the answer is written. Seer found 87% of SearchGPT citations matched Bing's top results, so an engine's index shapes which pages are eligible. Record the engine with every run, because a gap on one engine and not another points at retrieval before it points at content.
What should I do with gaps I cannot fix by publishing?
Route them, do not delete them. Third-party displacement belongs to whoever owns reviews and analyst relations, with the exact domains named from your citation record. Gaps where a rival honestly fits the question better belong to product and positioning. Both are real results, and recording them keeps the content queue finishable.



