ResourcesAugust 31, 2026 · 12 min read

Topic Clusters: What AI Engines Actually Reward

The authority story has no evidence behind it. Question coverage is what a cluster really buys.

Zach ChmaelLast updated August 31, 2026

TL;DR

Search the published literature for a study showing that pillar-and-spoke architecture makes an AI engine more likely to cite you, and you will not find one. The claim is everywhere in GEO advice and rests on nothing anyone has measured in public. Google's own documentation points the other way, stating that there are no additional requirements or special optimizations to appear in AI Overviews and AI Mode, and no new machine readable files or markup are needed. The defensible case for topic clusters is smaller and more useful: clustering is a planning discipline that produces question coverage and a consistent entity description, and while neither guarantees a citation, coverage at least addresses the queries a documented fan-out actually issues.

The reason to care is where buying now starts. 51% of software buyers now begin research inside an AI chatbot, G2 found AI chatbots are the single largest influence on B2B shortlists, and 6sense found buyers already know 3.8 of the roughly 5 vendors they evaluate before they speak to a seller. Owning a territory means being present when those questions get asked, one question at a time.

Do AI engines reward pillar pages and topic clusters?

No published study supports that, and the strongest evidence in the area argues against structural tricks generally. C-SEO Bench, published at NeurIPS 2025, tested ten conversational-SEO methods and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains, and the benchmark ran across two tasks and six domains rather than one narrow setup. Rewriting a page into a preferred shape mostly did nothing.

Google's guidance has the same shape. It asks for helpful, reliable, people-first content, and its AI features documentation states that AI Overviews and AI Mode use a query fan-out technique. Google's AI Mode announcement and updates from I/O 2025 cover the same feature set. Nothing in either says a hub page with spokes scores higher than a single page that answers the question well.

So drop the authority story. What survives is the part you can defend: a cluster is a checklist that stops you leaving buyer questions unanswered, and a style guide that keeps your name attached to the same description on every page in the set.

What does a cluster actually do for retrieval?

It changes what there is to retrieve. Google documents a query fan-out technique for its AI features, and OpenAI documents that ChatGPT search turns a conversation into one or more targeted queries, with its web search tool caching results for about 10 minutes before going back out. Each of those queries is answered by a passage on some page, which may or may not be yours. Neither company documents any preference for site architecture either way.

Position inside a document matters too. Liu et al. showed that models use long contexts unevenly, with relevant text buried in the middle used less reliably than text near the edges. A dedicated page that answers one question in its first screen is easier to use than the same answer at paragraph 40 of a 3,000-word overview.

The second input is identity, and here the evidence is a caution rather than a lever. An audit of 443 entity-oriented retrieval configurations on a TREC newswire test collection found that once entity selection was kept independent of the relevance judgments used for evaluation, none of its 437 unsupervised configurations beat a BM25 baseline of 0.292 MAP. Consistent naming is therefore a hygiene measure rather than a retrieval lever: it will not win you a passage, but three pages describing your company three different ways make the resolution harder and buy nothing for it.

How do you choose the territory worth covering?

Start from the decisions a buyer has to make before choosing, not from a keyword export. Write down the five to fifteen questions in your category that a real evaluation forces: what the category is for, what the alternatives are, what it costs, what breaks, what it connects to, who it is wrong for. That list is your territory, and it is finite in a way a keyword tree never is.

Then check who currently answers each one. Profound found 57% of AI citations point to sources brands do not control, so for some questions the answer already lives on a review site or a forum thread that a new page of yours will not displace. Seer found 87% of SearchGPT citations matched Bing's top results in a sample of 500, which tells you the contest for a passage is the existing top of search.

Where a third party owns the answer, the work is different in kind. AirOps examined 548,534 pages in its study of how ChatGPT chooses the sources it cites, and G2's research reports that one-third of buyers purchased from a vendor they had never heard of before the process began. A question whose answer sits on somebody else's property is not a page you can commission; it is presence you have to earn. Sorting the two kinds apart at this stage is what keeps a territory list from turning into a publishing plan.

An absence is a finding. It becomes an assignment only when the question is one you can answer with evidence and the answer is not already locked up by a third-party source you would be better off earning presence in. Some gaps describe a buyer you do not serve, and the right response there is no page at all.

How do you decide which questions deserve their own page?

Give a question its own page when a buyer would ask it separately and the answer needs evidence of its own. Consolidate when two questions share the same evidence and the same decision. That test replaces the older advice about whether a topic can support enough supporting posts to fill a cluster, which was a volume target wearing a strategy costume.

Whatever you publish has to carry proof. The GEO study published at KDD 2024 found that adding statistics, quotations and citations raised citation visibility in its benchmark while keyword stuffing did not, and scarcity and exclusivity framing measurably reduces how often a model recommends a product. Sales language on a cluster page is a cost, not a neutral choice.

Page count on its own is a weak bet, and the two things it is supposed to buy come apart under measurement. 38% of AI Overview citations rank in the organic top 10 across 4 million AI Overview URLs, down from roughly 76% in July 2025, so ranking and being cited are separating into different outcomes rather than one prize. Separately, the top-ranking page now sees a 58% lower average clickthrough rate when an AI Overview is present, so the pages that do rank return fewer visits than they used to. Publishing more spokes buys neither result by itself.

How do you keep your entity description consistent across a cluster?

Write the description once and reuse it without edits. One sentence for what the company is, one for the category it sits in, one for who it is built for, and identical product names in identical casing on every page in the territory. This is the cheapest part of the cluster idea to get right, and the only part that costs you something when you get it wrong.

Say the same facts in markup where markup is defined. Schema.org's getting started documentation and Google's explanation of structured data define types for organization, product and FAQ facts, and the types Google Search supports are published in full. That is what markup is documented to do. Google states no new markup is needed for its AI features, and no published study we could find shows markup changing what an AI engine cites.

Then confirm a machine can read any of it. Vercel's crawler research with MERJ documented that OpenAI's and Anthropic's crawlers do not render JavaScript, while Google's AI features inherit the same Web Rendering Service capabilities as Search. Cloudflare tested the 200,000 most visited domains for the things an agent needs to transact, including machine-readable structure and predictable access. A cluster rendered client side is an empty shell to every crawler in the first group.

How do you tell whether the territory is actually covered?

Measure appearance per question across repeated runs, never once. SparkToro tested how consistently AI engines recommend brands and found them highly inconsistent, and Search Engine Land's write-up of a separate study reports that AI recommendation lists rarely repeat exactly. One run of one cluster question tells you close to nothing about coverage.

A 2026 variance-components study found run-to-run noise large enough to swamp real differences in small samples, and work on quantifying uncertainty in AI visibility argues for confidence intervals instead of point readings. Neither paper sets a run count. Ten or more per question, spread across days and engines, is the floor we use before calling a question covered or missing.

Report it question by question. The IAB framework separates presence, prominence, portrayal and persuasion, which are different outcomes with different fixes. A territory with eight of twelve questions answered and four absent gives you a work list; a single composite score hides which four.

See the workflow: observed answers, useful drafts, human approval, and publication verification.

How does Trovance help you own a question territory?

Trovance treats the territory as a set of tracked questions rather than a content calendar. You define the questions your buyers ask before choosing, and the platform runs them repeatedly across AI engines, preserving each answer run as a snapshot that records who was mentioned, who was cited, who was recommended, and which sources carried the answer. Answer coverage is then reported per question, so you can see which parts of the territory are answered without you.

That record separates the two problems a cluster is supposed to solve. A question where competitors appear and you do not is a coverage gap, and the snapshots show whether the winning answer was built on the competitor's own page or on a third-party source you would have to earn. A question where the answer describes your company inaccurately is an entity problem, and your Brand Core holds the claims you are entitled to make, the proof behind each one, and the description that every page in the territory should reuse.

From there the work becomes specific. Recommended actions name the asset the evidence record says is missing, drafts are produced from approved claims rather than from a topic prompt, and a person reviews everything before it publishes. After a page goes live, the next analysis cycle reruns the same tracked questions so you can compare answers against the ones you preserved and see whether coverage moved.

Trovance will not tell you that a pillar-and-spoke structure causes citations, because no published evidence we could find supports that. It does not promise guaranteed citations, rankings or recommendations, it does not produce a universal visibility score, and it does not publish without human approval. What it offers is a durable record of what the engines said about your territory over time, and honest attribution of what changed after you acted.

What should you do this week?

Write the question list first, before any page. Twelve questions your buyers actually ask, in their words, with a note on who answers each one today. That single document does more for coverage than a diagram of hubs and spokes.

Then run each question ten times across at least two engines, on different days, and record appearance and sources. Fix the mechanical failures the runs expose: pages an AI crawler cannot read, product names that vary across the set, claims with no evidence attached. Only then commission new pages, one per question that survives the test in the third section above.

Be honest about pace. Retrieval fixes can show up within weeks, earned third-party presence takes months, and measurement variance means early readings will wobble either way. If you want the question list run continuously instead of by hand, start a free Trovance analysis and watch coverage question by question.

Plan the territory

Make each page usable

FAQs

Do topic clusters and pillar pages improve AI citations?

No published study shows that structure causes citations, and C-SEO Bench found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. Clusters still help, for a different reason: they force coverage of the questions buyers actually ask, and they force one consistent description of your company across the whole set.

What is a pillar page for if engines do not score it?

It works as a planning artifact and a navigation aid for readers. A pillar collects the questions in a territory, links to the page answering each one, and gives your team one view of what is still missing. Treat it as an index over real answers rather than as a ranking mechanism.

How many pages should a cluster contain?

As many as there are buyer questions that need evidence of their own. The older rule about filling a cluster with a fixed number of posts is a volume target, not a coverage test. Ahrefs measured a 58% lower average clickthrough rate for the top-ranking page when an AI Overview is present, so volume does not mean visits.

How do topic clusters and pillar pages affect entity clarity?

Building a cluster forces you to describe the company identically on every page: one category, one set of product names, one boilerplate sentence. An audit of 443 entity-oriented retrieval configurations found none of its 437 unsupervised setups beat a BM25 baseline, so treat consistency as hygiene rather than as a retrieval lever.

Should I write the pillar page or the cluster articles first?

Write the pages that answer specific buyer questions first, because those are what a passage-level retrieval step can use. Assemble the pillar afterward as an index over what exists. Publishing an overview before the answers exist leaves you with a page that mostly links to nothing and proves nothing.

How do I know whether a cluster is working?

SparkToro's research found AI engines highly inconsistent when recommending brands, so track appearance per question across ten or more runs on different days rather than reading a single answer. Report coverage question by question, and note whether the answer came from your page or from somewhere else.

Does structured data make my cluster more citable?

Google states there are no additional requirements or special optimizations to appear in its AI features, and no new markup is needed. Schema.org and Google's documentation define organization, product and FAQ types for Search appearance, and no published study shows markup changing what an AI engine cites.

Related resources

All field notes →