ResourcesSeptember 7, 2026 · 11 min read

AI Search Visibility Audit: Measure Retrieval, Citation, Comparison, and Recommendation

Replace one blended GEO score with a four-stage audit that shows where your brand disappears and what evidence-backed action should come next.

Zach ChmaelLast updated September 7, 2026

TL;DR

Audit AI search visibility as 4 separate stages: retrieval, citation, comparison, and recommendation. A brand can be findable without being cited, cited without surviving a comparison, and compared without entering the final recommendation. One blended score hides those differences and sends teams toward the wrong fix.

The practical move is to collect repeated answers for a defined buyer question, preserve the sources and wording, and classify what happened at each stage. A September 2026 preprint tested 50 buying questions across 6 engines and 15 runs per question-engine cell. In that sample, 1 run exposed only 62-77% of the brands found across 5 runs, while the median question surfaced 38 organizations across engines. Those figures are study results, not universal sampling rules, but they show why a single screenshot is a weak diagnosis.

Can the system retrieve usable evidence about your brand?

Retrieval asks whether the system can find material that matches the buyer's question well enough to use. Start by checking whether your brand appears, which pages or domains surface, and whether the retrieved evidence covers the relevant product, category, geography, use case, and current facts. Do not jump from absence to "write more content."

Retrieval is the eligibility layer of the audit. A technically accessible page can still be a poor match for the question. A strong category page can still omit the proof needed for a narrow buying criterion. A model may also answer from parametric knowledge without showing a source, which limits what an outside auditor can verify.

A 2026 system paper on inventory-grounded AI search gives this distinction a concrete form. Its example separates 3 explanations for a failed product-search result: the desired item is absent, retrieval misses an available match, or the system retrieves the item but final selection rejects it. That is one commercial implementation, not a map of every public engine, yet the diagnostic split is useful.

For a public audit, record 4 things for every run: the exact question, answer text, visible citations, and timestamp. Then add the brand pages that should plausibly satisfy the request. If the answer cites your category peers but never your relevant page, you have a retrieval hypothesis. If it cites your page but omits the brand from the answer, the failure sits later in the chain.

Treat hidden mechanics as unknown. Public outputs let you observe mentions and exposed sources, not prove the engine's internal retrieval path. I use the word "retrieval" here as an audit stage tied to available evidence, not as a claim that every model executed the same pipeline.

Did the answer cite the source that carries your claim?

Citation asks whether the answer exposes a source that supports the statement being made. Measure the cited domain, exact page, nearby claim, and whether the page actually contains the evidence. A citation count alone cannot tell you whether the source supports the answer, whether the brand was considered, or whether the citation changed the recommendation.

The repeated-query study makes the sampling problem visible. Across 50 buying questions, 6 engines, and 15 runs, cited-domain accumulation continued through every horizon the author tested. In 4 deeper web-search cells observed for 24 runs, the paper reports that cited domains were still being added at run 24. That means a short audit can under-sample the source pool even when the answer format looks consistent.

Repeated-query study Figure 1 shows brand accumulation across six AI engines

Source: Żatuchin, Sampling Completeness in Generative Search, Figure 1. Screenshot of the canonical arXiv PDF, accessed September 7, 2026.

The figure shows a second distinction: 5 engines answering without retrieval kept adding organizations through run 15, while the 1 retrieval-enabled engine leveled earlier in this sample. The study's method, engines, query bank, and extraction process bound that result. It supports repetition and source preservation, not a fixed 15-run standard for every audit.

Review citation quality in 3 passes. First, confirm the cited page is live and accessible. Second, locate the sentence, table, product detail, or evidence block that supports the answer.

Third, note any conditions the answer dropped, such as date, geography, inventory, denominator, or tested system. A citation is useful evidence only when its relationship to the claim survives inspection.

This also reveals a common repair. If your page is retrieved and cited for a category statement but lacks the specific comparison criterion the buyer asked about, the next move may be to strengthen that page. If an unrelated third party carries the needed proof, the job may involve authority and distribution rather than publishing another near-duplicate article.

Does your brand survive the comparison criteria?

Comparison asks whether the answer evaluates your brand against the buyer's actual constraints. Capture the alternatives named, the criteria applied, the evidence used for each criterion, and the language that moves a brand into or out of contention. A mention in a long list is different from a supported comparison against price, fit, availability, workflow, or risk.

This is where retrieval evidence becomes a decision structure. The inventory-grounded paper's architecture separates retrieval from selection and conditions both stages on current inventory evidence. Its online path builds an inventory portrait, gives relevant policy guidance to the retriever, and passes retrieved items into a selection stage. The published system is specific, but it demonstrates why "found" and "chosen" should not be treated as synonyms.

Inventory-grounded AI search Figure 1 separates grounding retrieval and selection

Source: Zhou et al., Inventory-Grounded Policy-Level Optimization, Figure 1. Screenshot of the canonical arXiv PDF, accessed September 7, 2026.

The figure also shows why current facts matter. The system probes inventory, constructs a scene profile, retrieves candidates, and applies selection guidance. The authors explicitly note that incomplete metadata, sparse coverage, or noisy retrieval probes can limit diagnosis. For marketers, the bounded implication is simple: stale or missing evidence can break the chain before the final answer is written.

Build a comparison record with 5 fields: buyer question, named alternative, criterion, supporting source, and observed judgment. Repeat that record across engines and runs. You are looking for patterns such as a missing criterion, stale description, ambiguous category fit, weak proof, or inconsistent naming.

Those are actionable evidence gaps. A count of appearances cannot explain them.

Keep the criteria buyer-led. If the prompt asks for a tool suited to a lean marketing team, capture implementation burden and workflow ownership. If it asks for a product available today, capture current availability.

Do not substitute the dimensions your company wants to win. The audit is only useful when it preserves the decision the buyer asked the system to make.

Does the brand enter the final recommendation?

Recommendation is the narrowest outcome: the brand appears in the answer's final choice set with an observable reason. Record its position, wording, qualifiers, alternatives, and cited support. Keep recommendation separate from mention, sentiment, and citation because each can move independently across repeated runs and across engines.

In the repeated-query preprint, the median question drew 38 organizations across 6 engines, and a median 15 organizations appeared in exactly 1 engine. The best single engine showed 83% of the cross-engine union. These numbers describe that study's sample, but they warn against treating one engine as the whole market or one answer as a stable shortlist.

Use 4 labels for each answer: recommended, compared but not recommended, mentioned without comparison, and absent. Add a fifth label, indeterminate, when the answer format does not support a clean judgment. This prevents an analyst from scoring every mention as success or interpreting every absence as a technical failure.

Then preserve the reason. "Best for enterprise controls" carries more diagnostic value than rank 2. It points toward the criterion and proof the system used.

If the reason is wrong, the work may be a factual repair. If it is accurate but weakly supported, the work may be better evidence. If the brand is a poor fit for that buyer question, the correct action may be no content change.

Recommendation remains an output observation. It does not establish preference across all buyers, cause a click, or prove revenue impact. Connect it to later behavior only with a separate measurement design. I would rather keep that boundary visible than turn an easy-to-count answer feature into an outcome it cannot carry.

How should you run a four-stage GEO audit?

Run a bounded test around 1 buyer decision, not a giant prompt panel. Define the question, engines, run count, location, account state, date, and classification rules before collection. Save every raw answer and source. Then classify retrieval, citation, comparison, and recommendation separately before deciding what work the evidence permits.

Use this 6-step sequence:

  1. Choose 1 buyer question. Write the decision and the context that makes it commercially meaningful.

  2. Freeze the test contract. Record engines, exact wording, settings, run cadence, and what counts at each of the 4 stages.

  3. Collect repeated outputs. Use enough repetition to expose variation, while refusing to call any number a universal optimum.

  4. Inspect every visible source. Confirm the page, passage, freshness, and claim relationship rather than counting domains only.

  5. Classify the first broken stage. Retrieval, citation, comparison, and recommendation imply different fixes.

  6. Choose 1 action and re-observe. Preserve a baseline, ship the approved change through your normal process, and test the same contract later.

A simple audit table can keep the evidence legible:

Run Retrieved evidence Citation support Comparison criterion Recommendation state
1 Relevant product page visible Claim supported Team size Compared, not selected
2 Third-party category page Partial support Implementation effort Recommended
3 No visible source Unverifiable None stated Mentioned only

Do not average away the mechanism too early. A 40% recommendation rate and a 40% citation rate can describe completely different answer sets. Keep row-level evidence until you can explain which stage changed and why. Aggregate metrics are summaries, not substitutes for the underlying answers and sources.

End with a governed action: refresh a canonical page, add missing proof, correct a stale fact, strengthen internal discovery, improve distribution, commission original evidence, or hold. The audit earns its value when the output becomes specific work with an owner and a later observation point.

See how Trovance carries an observed buyer question from answer evidence to a reviewable action.

How does Trovance turn the four-stage audit into action?

Trovance observes a buyer question across answers, citations, comparisons, and recommendations, then keeps the source evidence attached to what happened. Instead of stopping at a visibility score, your team can inspect where the brand first disappears and decide whether the next move is a technical repair, a clearer claim, stronger proof, a refreshed canonical page, an authority asset, or no content change.

The workflow turns that diagnosis into governed work. The question, answer, sources, claim boundaries, recommended action, and review state stay connected, so a writer or operator receives evidence rather than a vague instruction to "improve GEO." Human review remains at the decision points, and the original observation provides a baseline for the next measurement cycle.

That means fewer disconnected dashboards and fewer speculative briefs. You can move from "our score changed" to "this question retrieved the wrong evidence, the comparison used a stale criterion, and this is the specific asset we should repair." The mechanism helps a lean team choose, produce, review, and later re-observe one defensible action at a time.

Start a free visibility scan to inspect a real buyer question, see the sources shaping the answer, and turn the first evidence gap into accountable work.

FAQs

What is a GEO audit?

A GEO audit examines how a brand appears in AI-generated answers for defined buyer questions. A useful audit preserves each answer and its visible sources, then separates retrieval, citation, comparison, and recommendation. It should produce a bounded diagnosis and next action rather than one unexplained visibility score.

How many times should you repeat an AI search prompt?

There is no universal run count. Repeat enough to expose variation for your question, engine, settings, and time window, then report that method. One 2026 preprint used 15 runs per question-engine cell, but its design is evidence about that sample, not a required threshold for every audit.

Is an AI citation the same as a brand recommendation?

No. A citation identifies a visible source associated with an answer, while a recommendation places a brand in a choice set with some stated or implied reason. The cited page may support background information, a competitor, or one criterion. Inspect the exact claim-source relationship before classifying the outcome.

Can public testing prove that an AI system retrieved my page?

Public testing can preserve answers, links, citations, and repeated output patterns. It usually cannot expose every hidden retrieval, ranking, or generation step. Treat a visible citation as evidence that the source appeared in the response, and treat uncited internal processing as unknown unless the provider supplies stronger documentation.

What should I fix when my brand is absent?

First identify the earliest supported failure stage. Check access and relevance, then source support, comparison criteria, and final recommendation. The right action may be a technical repair, stronger proof, a refreshed canonical page, clearer positioning, distribution, or no change when the brand is not a fit for that buyer question.

Does better AI visibility prove more revenue?

No. Retrieval, citation, comparison, and recommendation are answer-level observations. Traffic, activation, pipeline, and revenue are later outcomes with separate denominators and attribution requirements. Preserve those stages rather than joining them with assumed causality; a visibility audit can prioritize work, while commercial impact needs its own measurement design.

How does Trovance support a GEO audit?

Trovance observes buyer questions and preserves the answers, citations, comparisons, recommendations, and sources needed for review. It helps a team identify the visible failure stage, connect that evidence to a specific content or authority action, keep human approval in the workflow, and establish a baseline for later observation.

Related resources

All field notes →