A Visibility Gap Is Not a Content Brief

A Visibility Gap Is Not a Content Brief

Here is the workflow almost every AI visibility tool encourages, whether or not it says so out loud.
Here is the workflow almost every AI visibility tool encourages, whether or not it says so out loud.

5 min

Zach Chmael

In This Article

Seven reasons your brand is missing from AI answers. Only three are content problems, and one of them means you should not publish anything at all.

Updated

A Visibility Gap Is Not a Content Brief

Here is the workflow almost every AI visibility tool encourages, whether or not it says so out loud.

You open the dashboard. Your brand is absent from six of twelve tracked prompts. The tool shows you the six, maybe highlights which competitors appeared instead, and offers to help you "close the gap." You take the list to your content team. Six briefs get written. Six articles get published. Ninety days later you check again, and the numbers have moved a little, or not at all, and nobody can say why.

The problem starts before the first brief. Absence from an answer is a symptom. It is not a diagnosis, and it is definitely not an assignment.

Is the gap even real?

Often, no. Which is the most uncomfortable place to start, so let's start there.

Rand Fishkin of SparkToro and Patrick O'Donnell of Gumshoe.ai ran what is currently the most useful study on this question. Six hundred volunteers ran 12 identical prompts through ChatGPT, Claude, and Google's AI nearly 3,000 times, and every response was normalized into an ordered list of brands so the lists could be compared.

The odds of getting the same list twice came in under 1 in 100. The odds of the same list in the same order were closer to 1 in 1,000. List length swung from two or three names to more than ten.

That is not a bug in the models. Large language models are probability engines, built to generate variation rather than return a stable ordered set. A 2026 variance-components study reached the same conclusion from the measurement side, finding brand-ranking reliability near 0.01 for a single answer and reporting that reliability comes from spreading measurement across models and phrasings rather than repeating one prompt. Search Engine Land's write-up puts the implication plainly: any product selling AI rank movement is selling fiction.

One thing did hold up. Visibility percentage across many runs was meaningful. Some brands showed up in 60% to 90% of responses for a given intent even as their position bounced around. Repeat presence means something. Exact rank does not.

So before you write a single brief, the first question is whether you measured a gap or measured noise. If your evidence is one run of one prompt on one engine, you have not found a problem. You have found a sample size of one.

This is fixable without buying anything: run each prompt several times, across more than one engine, and record what comes back. It is tedious, which is the honest reason most teams skip it. (It is also the first thing we automated in Trovance, because a diagnosis built on one run is worse than no diagnosis.)

This gets worse in large categories. The same research found that smaller markets produced more stable results, with niche B2B tools and regional providers clustering around a few familiar names, while big categories scattered into chaos. If you sell into a crowded space, your single-run readings are the least trustworthy of anyone's.

What are the seven reasons a real gap exists?

Once you have established a gap is persistent rather than probabilistic, it has a cause. In my experience there are seven, and they are not interchangeable.

1. Identity. The model cannot tell what you are. Your positioning is abstract, your category language is invented, or your own site describes you differently than your LinkedIn, your G2 profile, and your press coverage do. The model has conflicting information about a basic fact and resolves the conflict by naming someone clearer.

2. Relevance. The model knows what you are and does not think you fit this prompt. Sometimes it is wrong, and G2 found one-third of buyers purchased from a vendor they had never heard of after a chatbot named it, so unfamiliarity alone does not disqualify you. Sometimes it is right, which we will come back to.

3. Evidence. You claim outcomes but nothing verifiable supports them. The Princeton GEO study found citations, quotations, and statistics produced the largest visibility gains of nine tested tactics, up to 40%. Adjectives without numbers, case studies without specifics, "trusted by leading companies" without naming one.

4. Comparison. Someone asked which of several tools is better and you have never published an honest comparison. The model assembles one from your competitors' framing, because that is the only framing available.

5. Reputation. Your third-party footprint is thin. Review sites, industry publications, community discussions, and analyst mentions carry disproportionate weight, and you are barely present in them.

6. Freshness. Your public facts are stale. Pricing changed, the product shipped three major capabilities, the integration list doubled, and none of it is on a page a retrieval system can reach. 6sense found 58% of buyers contacted sellers earlier than they otherwise would have, specifically because vendor websites did not answer their questions about pricing, security, and AI capabilities.

7. Technical access. The content exists and cannot be read. This one is invisible from a dashboard and it is more common than anyone expects.

Which of these actually call for content?

Three. Maybe four.

Evidence, comparison, and freshness are genuine publishing problems. If nothing verifiable supports your claims, you write proof pages and case studies with real numbers. If no honest comparison exists, you publish one. If your facts are stale, you update the pages that carry them. In each case, the missing artifact is a document, and creating the document is the fix.

Identity is halfway. Some of it is a content fix, in that your pages should state plainly what you are. Most of it is a consistency problem across surfaces you do not fully control, which is closer to an operations job than an editorial one.

The other three are not content problems at all, and treating them as content problems is how teams burn a quarter.

Reputation is a distribution problem. Publishing more on your own domain does not build a third-party footprint. That work is earned media, review generation, community participation, and analyst relationships. Semrush's AI Visibility Index, built on 126 million U.S. AI search prompts across ChatGPT, Gemini, Google AI Mode, and AI Overviews, found that a company's AI narrative is no longer shaped solely by owned websites: platforms also draw on customer reviews, community discussions, independent publishers, retailers, and industry sources. Companies producing the most content are frequently still absent, because the missing signal was never on their site to begin with.

Technical access is an engineering problem. Vercel and MERJ's crawler analysis found GPTBot fetched JavaScript files in 11.5% of requests and ClaudeBot in 23.84%, and neither executed them. Their conclusion was that none of the major AI crawlers render JavaScript, with Gemini the exception because it inherits Google's infrastructure. If your comparison page is client-rendered, publishing a second one changes nothing. Both are invisible. You need server-side rendering, not another brief.

Relevance may not be a problem. More on this below, because it is the one everyone skips.

What if the model is right that you don't fit?

Then the correct action is to publish nothing.

This is the part no visibility vendor will tell you, and the reason is structural rather than dishonest: if a company sells content production, every gap has to look like a content gap. A diagnosis of "this prompt is not yours" produces no work order.

But some prompts genuinely are not yours. If you sell a $50,000 enterprise platform and you are absent from "cheapest project management tool for freelancers," you have not found a visibility problem. You have found evidence that the model understands your positioning correctly. Chasing that prompt means publishing content aimed at buyers you cannot serve profitably, attracting trials that will not convert, and diluting the entity signals that make you legible for the prompts that do matter.

The same logic applies to prompts describing an adjacent capability you do not have, a geography you do not serve, and a company size you cannot support.

There is a real cost to getting this wrong in the other direction, though, and it deserves saying. "The model is right that we don't fit" is also the most comfortable conclusion available, and comfortable conclusions need a higher evidence bar than uncomfortable ones. The test I use: would I want this buyer if they walked in tomorrow? If yes, it is a gap. If I hesitate, it was never mine.

How do you tell the seven apart?

By looking at the answer rather than the score.

Most dashboards throw away the thing you need. A percentage tells you that you were absent. The raw answer tells you why. If you keep the actual text and the citations behind it, the diagnosis is usually visible in a couple of minutes:

  • The model described your category incorrectly, or described you as something you are not. Identity.

  • The model described you accurately and chose someone else as a better fit. Relevance, and possibly correct.

  • Competitors were cited with specific numbers and you were cited with adjectives, or not cited at all. Evidence.

  • The answer was a comparison, and every source was a competitor's page or a third-party roundup. Comparison.

  • The citations were review sites, forums, and publications rather than vendor sites, and you were absent from all of them. Reputation.

  • The model stated something about you that used to be true. Freshness.

  • The model made no reference to a page you know exists and covers the topic well. Technical access, and go fetch that page as an agent would to confirm.

One caveat worth building into the habit: check more than one engine before concluding anything. Semrush found citation patterns differ sharply by platform, with the overlap between mentioned brands and cited domains as low as 30% on Gemini. A brand can be well represented on one engine and absent on another for reasons that have nothing to do with content quality.

None of this requires special tooling. It requires not discarding the evidence. You can do the whole thing with a spreadsheet and an afternoon, and for a while that is exactly what I did: paste the raw answers into a sheet, tag each one by cause, and watch the pattern emerge.

Trovance exists because I got tired of doing it by hand, but the method matters more than the tooling. Keep the answers. Read them. Sort by cause.

What does this change about how you work?

It moves the decision earlier, and it makes "no action" a legitimate output.

The dashboard workflow is: gap, brief, publish, hope. The alternative is: confirm persistence across runs, read the actual answer, name which of the seven it is, then choose the action that matches the cause. Sometimes that is a page. Sometimes it is a pricing update, a review campaign, an engineering ticket, or a conversation with product about a capability you genuinely lack.

That last category is worth dwelling on. Some visibility gaps are honest reports about your product. If the model consistently recommends competitors for a use case because they support something you do not, no amount of content fixes it. The gap is real, the diagnosis is accurate, and the correct owner is not marketing.

A system that can only produce content briefs will never tell you that.

Where this leaves the dashboard

Measurement is not the problem. Measurement that terminates in a percentage is.

A score can tell you that something is wrong. It cannot tell you whether the fix belongs to marketing, engineering, product, or nobody. The teams that will do well in AI search over the next two years are not the ones publishing the most in response to the most alerts. They are the ones who got good at telling the seven causes apart, and who developed the discipline to write nothing when nothing was the right answer.

That discipline is what we built Trovance around: preserve the real answers across engines and repeated runs, name the cause, and act only where action is warranted. Sometimes it recommends a page. Sometimes it recommends a review campaign, a pricing update, or an engineering ticket. Sometimes it says the prompt was never yours. That last one is a strange feature to ship, and we can ship it because we are not venture-backed and nobody is asking us to manufacture urgency to justify a subscription.

See how AI sees your brand


FAQs

Why is my brand missing from AI answers?

There are seven common causes: unclear identity, genuine irrelevance to the prompt, missing verifiable evidence, no published comparison, thin third-party reputation, stale public facts, and technical inaccessibility. Only three or four are solved by publishing content. The rest require product, distribution, or engineering fixes.

Is a single AI answer reliable enough to act on?

No. SparkToro and Gumshoe.ai found the odds of getting the same brand list twice from identical prompts were under 1 in 100, and matching order closer to 1 in 1,000. Persistent absence across many repeated runs is meaningful evidence. One run is a sample size of one.

Should I track my position in AI answers?

Position tracking is largely noise. The SparkToro research concluded ranking positions in AI answers are unstable enough to be effectively meaningless. Track how often your brand appears across many prompts run many times instead, since repeat presence proved to be a durable signal even when rank was not.

When should I not create content for a visibility gap?

When the model is correctly excluding you. If a prompt describes a buyer, budget, geography, or capability you cannot serve well, appearing there attracts unqualified demand and dilutes your entity signals. The test is whether you would want that buyer if they arrived tomorrow.

Can publishing more content fix a reputation gap?

No. Reputation gaps come from a thin third-party footprint across review sites, publications, forums, and analyst coverage. Content on your own domain does not create third-party signals. That requires earned media, review generation, and community presence rather than additional owned pages.

How do I know if a gap is a technical problem?

Fetch the page as an AI crawler would and see what comes back. Vercel and MERJ found GPTBot fetched JavaScript in 11.5% of requests and ClaudeBot in 23.84%, with neither executing it. If your content only appears after client-side rendering, most AI crawlers see an empty shell regardless of quality.

Do AI engines agree with each other about brands?

Often not. Semrush's index of 126 million prompts found citation patterns differ substantially across ChatGPT, Gemini, Google AI Mode, and AI Overviews, with mention-to-citation overlap as low as 30% on Gemini. Per-engine measurement is necessary because presence on one platform does not predict presence on another.

What should I do instead of writing a brief for every gap?

Confirm the gap persists across repeated runs, read the raw answer and its citations, identify which of the seven causes applies, then choose the matching action. That may be a page, a pricing update, a review campaign, an engineering ticket, or a deliberate decision to do nothing.

Be the answer.

Built to win the agentic web. Made to improve the human world.

Be the answer.

Built to win the agentic web. Made to improve the human world.

Be the answer.

Built to win the agentic web. Made to improve the human world.