ResourcesSeptember 4, 2026 · 12 min read

Agentic Marketing Hype: What to Actually Ignore

The claims that sell easiest have the weakest evidence, and the work that holds up is unglamorous.

Zach ChmaelLast updated September 4, 2026

TL;DR

The most-cited public test of conversational-SEO rewriting, C-SEO Bench, found only 3 of 54 unilateral conditions produced a statistically significant citation-rank gain. The agentic marketing advice we see sold hardest this quarter rests on a record about that thin. Google says in its own documentation that no new machine readable files or markup are needed to appear in its AI features. This piece takes the claims one at a time and says what the record supports.

The underlying shift is real, which is exactly why the weak advice sells. G2 found AI chatbots are now the single largest influence on B2B software shortlists, 51% of software buyers now begin research inside an AI chatbot, and organizational AI use climbed from 55% to 78% in a single year. A real shift with an immature evidence base is the condition under which confident tactics outrun the data.

Should you rewrite your pages in a more authoritative voice?

Not for the reason you are being sold. C-SEO Bench at NeurIPS 2025 tested conversational-SEO methods across two tasks and six domains and reported that only 3 of 54 unilateral conditions produced a statistically significant improvement in citation rank. Rewriting for tone, authority, or conversational cadence is the intervention we see sold most often in this category, and it carries the weakest supporting record in it.

The distinction matters, because the earlier result everyone quotes is real. The GEO study published at KDD 2024 reported gains from adding statistics, quotations, and citations, and its proceedings entry is easy to check for yourself. C-SEO Bench later evaluated that family of methods at larger scale and did not reproduce most of those gains, which is the reason to read the GEO result as a starting point and hold the question open.

Some persuasion language performs worse than neutral prose. Scarcity and exclusivity framing measurably reduces how often a model recommends a product, tested across 10 fictitious products. The C-SEO authors also report that traditional retrieval-side improvements outperformed the rewrite methods they tested. If a vendor's core offer is a voice rewrite, ask which of the 54 conditions they expect to move.

Do agents need a special machine readable file on your site?

For Google's AI features, the answer in Google's own documentation is no. Its page on AI features states that there are no extra requirements or special optimizations to appear, that no new machine readable files or markup are needed, and that pages must be indexed and snippet-eligible. That is the operator with the most to gain from a new standard telling you not to build one.

The same page describes a query fan-out technique, where one buyer question becomes several underlying searches. That is a reason to cover the sub-questions a buyer actually asks. It is not a reason to invent a file format. Google's guidance on succeeding in AI search repeats the ordinary advice: be indexable, be useful, be technically sound.

Structured data is still worth doing on its documented terms. Google explains what structured data markup is used for and publishes the specific types Search supports, with the schema.org getting started guide underneath it. Markup that describes what a page already contains is maintenance. Markup invented so agents will notice you is a proposal no engine has published a consumer for, and nobody has published a test showing whether it changes citations either way.

Is one dashboard reading a measurement?

No, and treating it as one is the most expensive habit in the category. SparkToro's research found AI engines are highly inconsistent when recommending brands, with the same prompt returning different vendor lists across runs, and Search Engine Land's write-up of a separate study reached the same conclusion about repeat lists.

The variance is measurable and it is large. A 2026 variance-components study found run-to-run noise big enough to swallow real differences at small sample sizes, and Ronald Sielinski's work on quantifying uncertainty in AI visibility argues the point in confidence intervals. Even output length wobbles: one reported measurement came to 188 tokens with a 56.26-token standard deviation, and that work reports the platform medians separately, keeping each platform's own figure visible.

The working rule is a base rate. Ask the same buyer question ten or more times, across days and across engines, and record how often each vendor appears. A screenshot of one answer is a sample, whichever way it went. A tool that shows a number moving week over week without showing the sample size behind it is reporting noise with a decimal point on it.

Is crawler traffic an audience?

A fetch is a request, and a request is not a reader. Bot volume is now a large share of what arrives at a site: Cloudflare's agentic internet bot report attributes 52% of crawler requests to AI-related crawlers, and its earlier work tracked GPTBot's share of the crawl mix climbing from 2.2% to 7.7%. Those are requests. None of them is a person deciding anything.

The gap between crawl and consequence is where this particular hype lives. Vercel's crawler research with MERJ documented that these crawlers largely read raw HTML, so heavy fetching says something about your server and very little about your standing inside an answer. Crawl logs can tell you whether you are reachable. They cannot tell you whether you were cited, recommended, or shortlisted.

Report the two things separately. Keep fetch data as a reachability signal, and keep key event rate, sessions with a key event divided by total sessions, as the signal for the humans who do arrive. The longer version of this argument sits in a crawl is not an audience.

Should you buy a single AI visibility score?

No, because no such measurement exists to buy. The IAB's work on measuring visibility in the AI era splits the problem into Presence, Prominence, Portrayal, and Persuasion and organizes its recommendations into 4 groups for brands. Those are four different questions with four different failure modes, and averaging them into one number destroys the only information you needed.

Industry indices are useful as market description and misleading as a personal scoreboard. Semrush's AI Visibility Index analyzed 126 million AI search prompts and is published as market research. It is not your score.

The vendors worth questioning are the ones who take that shape of data and hand you one number for your brand, in place of your appearance rate on the eleven questions your buyers actually ask. Measurement infrastructure that takes itself seriously publishes its own uncertainty: Cloudflare Radar documents 5 confidence levels and 6 normalization states so a reader knows what a number is worth.

Ask any scoring vendor which prompts, how many runs, and which engine version produced the figure. If those answers are proprietary, the score is a brand asset for the vendor rather than a measurement for you. We argued this at length in there is no such thing as an AI visibility score.

What does the published record actually support?

Four things, none of them exciting. The first is being fetchable and retrievable. AI crawlers largely do not execute JavaScript, and Cloudflare's agent-readiness work across the 200,000 most visited domains measures how much of the web an agent can actually read. A page an engine cannot read cannot be cited under any tactic you buy.

The second is evidence a passage can carry. That is the durable half of the GEO study: numbers, named sources, and quotable specifics travel into answers when adjectives do not. C-SEO Bench tested this family of edits at larger scale and found most conditions non-significant, so treat added evidence as a defensible default whose citation lift is still unproven.

The third is entity clarity. One body of research spans 443 entity-oriented retrieval configurations. Our own reading of that work is simple: a company that appears under three names gives a retriever an ambiguous string to resolve.

The fourth is the least comfortable: sources you do not own. Profound found 57% of AI citations point to sources outside the brand's control, Seer found 87% of SearchGPT citations matched Bing's top results in a 500-citation sample, and Ahrefs measured 38% of AI Overview citations ranking in the organic top 10 across 4 million AI Overview URLs from 863,000 SERPs in March 2026, down from roughly 76% in July 2025. That decline is its own warning against treating any ratio here as fixed.

See the workflow: observed answers, useful drafts, human approval, and publication verification.

How does Trovance tell a real signal from a sold one?

Trovance is built around the base rate this article keeps returning to. You define the questions your buyers actually ask, and those tracked questions are run repeatedly across AI engines. Each answer run is preserved as a snapshot with its context intact: who was mentioned, who was cited, who was recommended, and which sources carried the answer. Answer coverage is reported as an appearance rate over many runs, so one lucky answer never reads as progress.

That record is what lets you refuse a tactic on evidence. If your pages are absent from every snapshot while a raw fetch of them returns an empty shell, the diagnosis is retrieval, and no rewrite will change it, because these crawlers largely read the raw HTML they are served. If you appear but a competitor's numbers get quoted instead of yours, the diagnosis is proof. If the answers are assembled entirely from third-party sources that omit you, the work sits off your own domain, and publishing more of your pages is the wrong instrument.

Your Brand Core holds the claims you are entitled to make and the proof behind each one, so drafts are produced from claims you have already approved. Recommended actions name the specific missing asset the evidence record points to: the benchmark that would answer a competitor's quoted study, the comparison page a third-party source will never write for you. A person reviews everything before it publishes, and that review is a design decision we are not apologizing for.

What Trovance will not sell you is a single visibility score, control over what a model says, or a guaranteed citation. The engines are probabilistic, which is what SparkToro's research on inconsistent recommenders measured, and the sources shift. What the analysis cycle does instead is rerun the same tracked questions after your work ships and compare the new snapshots against the preserved ones, so you can see whether the answer moved or whether only the dashboard did.

What should you do this week?

Start by deleting work. Cancel anything on the roadmap that exists because agents supposedly need a new file, since Google's documentation already says no such file is needed for its AI features. Pause voice-and-tone rewrites justified by citation gains until someone can name the condition they expect to move.

Then spend the recovered hours on four checks. Fetch your five most important pages with JavaScript disabled and read what comes back. Write down the ten questions a buyer asks before choosing in your category, run each one ten times across two engines, and record appearance rates.

The other two checks are about proof. Audit whether your strongest claims have a number or a source attached to them. List the third-party pages your engines actually cite, and see which of them omit you.

Be honest about timelines while you do it. Retrieval fixes can surface in weeks, evidence improvements follow re-crawling and re-retrieval, and earning third-party coverage takes months. Nobody can promise a recommendation on any schedule. If you want those base rates without running them by hand, start a free Trovance analysis and let the snapshots tell you which of these claims applies to your company.

What to ignore

What holds up

FAQs

What agentic marketing trends should my startup ignore?

Four claims have weak published support: rewriting pages in a more authoritative voice, publishing a special machine readable file for agents, treating one dashboard reading as a measurement, and counting crawler hits as audience. A fifth, the universal AI visibility score, describes a measurement that no published methodology actually supports.

Does rewriting content in a conversational voice improve AI citations?

Rarely on its own. C-SEO Bench tested conversational-SEO methods across two tasks and six domains, and found only 3 of 54 unilateral conditions produced a statistically significant citation-rank gain. The earlier GEO study's improvements came from adding statistics, quotations, and citations, which is a change in evidence rather than a change in tone.

Do I need a special machine readable file so AI agents can read my site?

Google's documentation states that no new machine readable files or markup are needed for its AI features, and that pages must be indexed and snippet-eligible. No published test shows whether such a file helps or hurts with other engines, so treat that as an open question. Ordinary structured data, on Google's documented types, stays worth maintaining.

Is AI crawler traffic a sign of buyer demand?

No. A fetch is a request, and a request is not a reader. Cloudflare's reporting attributes 52% of crawler requests to AI-related crawlers, which describes infrastructure load on your servers. Track fetches as a reachability signal, and track key event rate separately for the humans who actually arrive on the page.

How many times should I run a prompt before trusting the result?

Ten or more, spread across several days and at least two engines. SparkToro found AI engines are highly inconsistent recommenders, and a 2026 variance-components study found run-to-run noise large enough to swamp real differences in small samples. Record appearance rates over runs instead of collecting screenshots.

Can any tool give me one AI visibility score?

Not honestly. The IAB splits the problem into Presence, Prominence, Portrayal, and Persuasion, which are four questions with different failure modes. Averaging them hides the one you are losing. Ask any scoring vendor which prompts, how many runs, and which engine version produced the number they sold you.

What should a startup do instead of chasing agentic marketing hype?

Four things with published support: make pages fetchable without JavaScript, attach numbers and sources to your claims, keep entity naming consistent everywhere, and earn presence in third-party sources you do not control. Profound found 57% of AI citations point to sources outside the brand's own control.

Related resources

All field notes →