ResourcesAugust 31, 2026 · 9 min read

Why We Built Trovance: The System After the Score

A score can show absence. It cannot decide whether the next move is content, product clarity, technical repair, third-party proof, or no action.

Zach ChmaelLast updated August 31, 2026

TL;DR

  • 📊 IAB separates AI visibility into 4 metric classes, which should not be flattened into one implied business outcome.
  • 🔬 Its 37-page framework classifies programs with fewer than 50 queries as exploratory, not decision-grade.
  • 🎲 A 200-query SearchGPT example produced apparent citation shares of 9.5% and 6.0%, but their 95% confidence intervals overlapped.
  • 🧪 C-SEO Bench found significant positive gains in only 3 of 54 unilateral rewrite cases under its corrected tests.
  • 🔁 Trovance uses a published 9-step product loop to connect observation to diagnosis, action, approval, verification, and learning.

An AI visibility score can tell you that your company appeared less often than a competitor. It cannot tell you why, whether the difference is stable, or what your team should do on Monday morning.

That missing decision is why we built Trovance.

Visibility tools made an important problem measurable. They showed companies that buyers can now meet an AI-generated explanation, comparison, or shortlist before reaching a website. But the dashboard usually ends at the exact moment the expensive work begins: deciding whether the result points to weak evidence, unclear positioning, stale product facts, inaccessible pages, thin third-party proof, a real product gap, or honest bad fit.

Trovance is the system after the score. It begins with the observed answer and ends with an accountable action, a published or routed artifact, and a record the team can inspect.

What does an AI visibility score actually measure?

An AI visibility score summarizes selected observations under one provider's prompt set, engines, sampling method, and weighting rules. It does not represent one universal market reality.

The industry's own measurement vocabulary is becoming more specific. IAB's August 2026 framework separates visibility into Presence, Prominence, Portrayal, and Persuasion. Presence asks whether the brand appears. Prominence asks where.

Portrayal asks how accurately or favorably the brand is described. Persuasion covers recommendation strength and post-citation behavior.

Those are 4 different measurement classes, not four interchangeable ingredients in a single score. A company can gain mentions while losing citation quality. It can rank higher in an answer that describes the product incorrectly. It can receive more citations without entering a buyer's shortlist.

IAB framework separates AI visibility into presence, prominence, portrayal, and persuasion metrics

Source: IAB, Measuring Visibility in the AI Era, page 4.

IAB also proposes 2 quality tiers: directional and decision-grade. Its criteria call fewer than 50 queries exploratory. Directional coverage includes at least 2 intent types, while decision-grade coverage includes all 4 named intent types and segments results by intent.

That framework does not certify Trovance or any other vendor. It does establish a useful refusal: do not treat every visible number as fit for every decision.

Why can the same score point to different problems?

The same low score can come from several different breaks, and each break has a different owner.

A company may be absent because the model could not retrieve an accessible page. The page may be easy to retrieve but lack the proof required for a specific buyer question. The proof may exist while the positioning fails to connect it to the buyer's job.

The product may not support the requested use case. The answer may simply be correct that another vendor fits better.

Observed result What to inspect Accountable action
Brand is absent from the answer Crawlability, retrieval, entity clarity, source set Repair access or canonical facts before drafting
Brand is mentioned but not cited First-party proof and third-party source coverage Strengthen the evidence path
Brand is cited but described incorrectly Product facts, pricing, integrations, positioning Correct the accountable source
Competitor is recommended instead Buyer criteria, comparative proof, legitimate fit Build a fair comparison or accept bad fit
Score changed without a clear reason Prompt set, engine mix, repeated runs, methodology Re-observe before acting
Gap is real but no owned page can answer it Existing inventory, proof permission, buyer job Produce the missing proof-backed asset

The score alone cannot choose among those routes. It compresses the symptom and hides the diagnosis.

Why should repeated runs come before a production brief?

Repeated runs should come first because the apparent winner in one sample may sit inside ordinary answer variance.

A 2026 arXiv preprint studied 3 platforms and 3 consumer topics. It collected daily samples over 9 days and ran a second regime at 10-minute intervals. The authors measured citation-share uncertainty and how similar the cited-domain sets remained across repeated versions of the same query.

In one worked example, 200 SearchGPT running-gear queries produced apparent citation shares of about 9.5% and 6.0%. The corresponding 95% bootstrap intervals were approximately 5.5%–12.5% and 4.0%–8.0%. They overlapped.

A three-and-a-half-point difference looked decisive until uncertainty was included. Across the tested settings, apparent gaps below roughly 5–7 percentage points often had overlapping intervals.

The paper is a preprint affiliated with a measurement provider. Its ranges are not universal constants. The useful mechanism is narrower: a single answer or score can look more precise than the underlying system is.

Why is publishing more content often the wrong response?

Publishing more content is wrong when the actual break is retrieval, product truth, positioning, third-party reputation, or legitimate bad fit.

C-SEO Bench gives this distinction teeth. The NeurIPS 2025 benchmark tested 2 tasks across 6 domains, using more than 1,900 queries and 16,000 documents. It compared conversational rewrites with retrieval-side SEO changes in controlled answer contexts.

Only 3 of 54 unilateral rewrite cases showed statistically significant positive gains under the corrected tests. In the GPT-4o-mini comparison, the first context position beat the tested rewrites in all 6 domains. In the multi-actor simulation, the best retrieval-side approach produced an AUC of 8.60 versus 1.88 for the best conversational-rewrite approach.

C-SEO Bench compares retrieval-side SEO with conversational content rewrites across six domains

Source: Puerto et al., C-SEO Bench, Figure 1.

The benchmark does not reproduce live Google, ChatGPT, Perplexity, Claude, or Gemini systems. It measures citation rank with candidate documents supplied to answer models. It does support a practical warning: better wording cannot repair every selection problem.

That is why Trovance does not convert each red indicator into a content assignment.

What should happen after a score finds a gap?

After a score finds a gap, the team should preserve the underlying answer, diagnose the break, assign the right owner, and verify the resulting action.

The operating loop has 9 steps:

Step Decision Receipt
Observe What did the system say, cite, compare, or recommend? Raw answer, prompt, engine, sources, time
Diagnose Is the break evidence, positioning, retrieval, product, reputation, or fit? Named gap with supporting records
Decide Which action is justified, and which is not? Ranked action, owner, reason
Produce or route Does this need an asset or a different team? Draft or routed task
Approve Are the claim, evidence, permission, and risk acceptable? Human decision and exceptions
Publish or execute What changed in the public or operating system? URL, object, or action receipt
Verify Did the intended state actually change? Live-state check
Re-observe How does the market answer now? Comparable rerun
Learn What can the team carry into the next cycle? Bounded finding with caveats

The sequence matters. Production before diagnosis makes the system efficient at solving the wrong problem. Measurement without action leaves a marketing team staring at a number it cannot use.

When is no action the right answer?

No action is right when the answer is accurate, the buyer is a poor fit, the evidence is insufficient to justify a public claim, or the proposed content would duplicate a stronger existing page.

This morning offered a small example inside our own editorial system. The campaign selector said the Trovance comparison with AirOps and Jasper was unfinished. The live article, generated route, sitemap, and publication receipt all showed that it had been published and verified on August 28.

The wrong response was to produce the comparison again. The correct response was to reconcile the stale publication ledger, run the selector tests, and let the normal editorial system choose a new topic.

That is not glamorous, but it is the product belief in miniature. A visible gap can be caused by stale state. More content would have created a duplicate while leaving the underlying break untouched.

See the workflow: observed answers, useful drafts, human approval, and publication verification.

What does Trovance do after the score?

Trovance connects the observed AI answer to the evidence decision, accountable action, and verifiable result.

A team begins with a real buyer question and the answers, citations, comparisons, and recommendations that follow. Trovance helps diagnose whether the break sits in evidence, positioning, product truth, retrieval, third-party reputation, or fit. It can then help produce the proof-backed asset or route the issue to the person who owns it.

Humans retain truth, permission, taste, spend, and publication. Trovance does not guarantee a citation or make every company the correct recommendation. It makes the work after observation visible enough to review: what the system said, why the team acted, what changed, and what the comparable rerun showed.

That is why a team would use Trovance instead of stopping at a dashboard. The score identifies a place to investigate. The product carries the investigation into governed work. Scan your AI visibility and inspect what should happen next.

FAQs

What is an AI visibility score?

An AI visibility score is a provider-specific summary of observations such as mentions, citations, answer position, portrayal, or recommendation behavior across a defined prompt set and engine mix. Its meaning depends on the prompts, sampling, repeated runs, weighting, and time window used to produce it.

Why is one AI visibility score not enough?

One score can hide disagreements among engines, buyer intents, answer types, and repeated runs. A company may gain mentions while losing citation quality or be described prominently but incorrectly. Keep the underlying answers and denominators available so the team can diagnose the movement before acting.

Should every visibility gap become a content brief?

No. A gap may come from inaccessible pages, stale product data, unclear positioning, weak third-party proof, a real product limitation, or honest bad fit. Diagnose the cause first. New content is justified only when a distinct buyer job and supportable evidence are actually missing.

How many prompts should an AI visibility program track?

There is no universal count. IAB labels fewer than 50 queries exploratory in its framework, but decision fitness also depends on intent coverage, repeated responses, platform separation, category variability, and the decision being made. A smaller directional monitor and a budget-setting program need different evidence.

Can better content improve AI visibility?

It can under some conditions, but no writing tactic works universally. Controlled GEO studies report gains for selected transformations, while C-SEO Bench found many rewrite effects weak or negative and larger effects from retrieval-side relevance. Treat content as one possible intervention after diagnosis, not the automatic cure.

How do you connect AI visibility to revenue?

Start with separate events: answer exposure, citation, referral, product action, qualification, pipeline, and revenue. Preserve the cohort, denominator, and observation window at every step. A visibility change can contribute to a business result without proving that one page or one model answer caused it.

Does Trovance guarantee citations or recommendations?

No. Trovance observes how AI systems explain, cite, compare, and recommend a company, then helps diagnose the evidence gap and produce or route the justified action. Models, source competition, buyer fit, and time still shape outcomes. Human review remains responsible for public claims and publication.

Scan my AI visibility

Related resources

All field notes →