TL;DR
📄 Google names exactly two conditions for appearing in AI Overviews and AI Mode: indexed, and eligible to show a snippet. No new files, no special markup, no AI-only format.
🧩 38% of AI Overview citations rank in the organic top 10, down from roughly 76% in July 2025, across 4 million AI Overview URLs and 863,000 SERPs.
🔍 87% of SearchGPT citations, in a sample of 500, matched Bing's top results. Overlap that high is consistent with a shared candidate pool, and the pipeline is not published.
🌐 57% of AI citations point to sources outside the brand's own site, so a plan that covers only your own domain covers part of the problem.
🧪 C-SEO Bench found only 3 of 54 unilateral conditions produced significant citation-rank gains, across two tasks and six domains.
Every AI Mode format sold to you as a requirement is sold against Google's own documentation, which names two conditions and no third. Google's documentation states that AI Overviews and AI Mode use a query fan-out technique, issuing multiple related searches across subtopics and data sources, and that there are no additional requirements to appear: no new machine-readable files, no special markup, no AI-only optimization. The entry condition is that a page is indexed and eligible to show a snippet.
So the short answer is this. AI Mode splits a buyer's question into sub-questions, searches for each one, and assembles a response out of what those searches return. Google described the same behavior as it expanded AI Mode after I/O 2025, and its guidance on succeeding in AI search repeats ordinary Search advice instead of naming a new discipline. Google does not publish what its retrieval operates on. What the documentation does say is that one question becomes many searches, which means the query set you are competing on is not your keyword list.
The outside measurements are consistent with that. Ahrefs found 38% of AI Overview citations rank in the organic top 10, down from roughly 76% in July 2025, across 4 million AI Overview URLs and 863,000 SERPs measured in March 2026. Whatever is driving it, citation position and organic position have come apart. The same team measured a 58% lower average clickthrough rate for the top-ranking page when an AI Overview appears, so the click you lose and the citation you might earn run on separate mechanisms.
What does Google say about how AI Mode picks sources?
Two conditions, and no third one. The AI features documentation names indexing and snippet eligibility as what a page needs, then points the rest of the way at ordinary Search practice. There is no application process and no separate submission route Google documents. The crawler documentation lists the user agents and their scopes, and that is the whole of the published surface.
Structured data does not change the entry condition. Google's introduction to structured data describes markup as a way to state a page's content explicitly for supported features, and the gallery of supported types is the finite list of what it can produce. No AI Mode type appears in it. Markup helps Google parse a page and opens no door labelled AI Mode.
The rest of the published guidance is the boring kind, which is the point. Google's helpful, reliable, people-first content guidance is what its AI-search advice points back to, and cloaking remains a violation under the spam policies, which forecloses serving one page to the fan-out and another to a reader. The rules are old and the retrieval behavior is new.
If the engine fans out, what is competing?
Sub-questions, and you cannot see them. The documented mechanism stops at the query set: one buyer question becomes many searches, and each of those searches has its own answer to find. A page earns its place by holding text that settles one of them cleanly, and you are working out which ones by inference rather than by observation. Research on multi-hop question answering across 1,186 chemistry questions shows the shape of that problem in a lab setting, where a question needing several facts is served by several retrievals combined rather than by one document that happened to cover everything.
Chunk size is where the lab evidence stops transferring, and it is worth being exact about why. In that academic retrieval system, the retrieved units averaged about 188 tokens with a 56.26-token standard deviation. That is one research group's pipeline, not Google's, and Google publishes no equivalent number. Read it as evidence that retrieval can operate below the document level somewhere. It is no guide to how long your own pages should be.
Entity resolution sits underneath all of it. Entity-oriented retrieval research across 443 configurations shows how much retrieval quality depends on the system working out what a name refers to. A paragraph that says "the platform" where it should say your product name is a paragraph a retrieval system has trouble attaching to a company.
Which sources does the fan-out reach?
Far more of the web than your own domain, which is why an owned-content plan answers only part of the question. Profound's citation research found 57% of AI citations point to sources outside the brand's own site. If most of what an answer rests on was written by someone else, earning presence in the sources already carrying your category can matter more than another post on your own blog.
Other engines run comparable machinery on comparable indexes. Seer found 87% of SearchGPT citations, in a sample of 500, matched Bing's top results. Overlap that high is consistent with a shared candidate pool, but the sampling is external and the pipeline is not published, so treat it as an observation about outputs.
The demand side has moved far enough to make the source pool worth budgeting for. 51% of software buyers now begin research inside an AI chatbot, and one-third of buyers purchased from a vendor they had never heard of before that research started. Whoever the answer cites is doing the introducing.
Which controls over AI Mode do you actually have?
Only subtractive ones. The only lever the documentation describes is snippet eligibility itself, and removing it removes you. No directive requests inclusion and no markup asks for it. The closest thing to an inclusion lever is being indexed and snippet-eligible in the first place, which is exactly the condition the documentation states.
That asymmetry is the honest summary of the mechanism. You can make yourself less available to a fan-out answer with a single line of configuration, and you can do nothing equivalent in the other direction. Anyone selling you the other direction is selling something Google has not described.
Rendering is the one place Google differs from the rest. Vercel's crawler research with MERJ documented that GPTBot and its peers read raw HTML without executing JavaScript, while Google applies the same Web Rendering Service capabilities throughout indexing. A client-rendered page can therefore be indexed by Google and remain an empty shell to another engine. Cloudflare's agent-readiness work covers the 200,000 most visited domains, and the question it asks is the one to ask of your own pages: does the answer arrive in HTML, without a browser. One fetchability test does not settle both questions.
Does rewriting pages for AI Mode work?
Mostly no, and the strongest test of the idea says so plainly. C-SEO Bench (NeurIPS 2025) evaluated conversational-SEO methods and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. The benchmark ran across two tasks and six domains, so the null result is not a narrow one. Rewriting for the engine is the tactic most often sold and the one with the least support behind it.
The GEO study (Aggarwal et al., KDD 2024) is the earlier result most rewrite advice still rests on, and it reported gains for adding statistics, quotations and citations. C-SEO Bench re-tested that family of methods across two tasks and six domains and did not reproduce them. Write with evidence because a reader needs it and because Google's published guidance asks for it. A benchmark promise is a weak reason to change a page.
That leaves the instruction Google's documentation already gives, reached from the opposite direction. A paragraph that settles a sub-question with a specific number and a named source is a paragraph a reader can use, and the retrieval question takes care of itself or it does not.
How do you audit your own fan-out coverage?
Start from the buyer's question and decompose it by hand. Write the five to ten questions a real buyer asks before choosing in your category, then split each one into the sub-questions an answer would have to settle: what it costs, what it connects to, what the alternatives are, what proof exists. That list approximates the fan-out far better than a keyword export does.
Then check coverage passage by passage. For each sub-question, find the specific paragraph on your site that answers it in place, without depending on the section above it for context. A missing answer is a content gap. An answer buried three sentences into a paragraph that opens with a transition is a retrieval gap, and retrieval gaps are cheaper to close.
Run each question at least ten times across several days before treating any single answer as a reading. SparkToro's research on consistency in AI brand recommendations, reported in Search Engine Land's write-up of it, and a separate 2026 variance-components study are the reason to bother. None of them hands you a number you can carry into your own category, which is exactly why you measure your own repeatedly.
Record what you find in named rungs rather than one number. The IAB's measurement work separates presence, prominence, portrayal and persuasion, which stops a passing mention from being logged as a recommendation. Being named in an answer and being the source the answer rests on are different outcomes with different fixes.
How does Trovance observe what AI Mode cites?
Trovance runs the audit above as a standing system instead of an afternoon of manual sessions. You define the market questions your buyers ask, and it runs them repeatedly across AI engines, preserving each answer run as an answer snapshot with its full context: who was mentioned and who was cited. Because snapshots accumulate, the record shows appearance rates across runs rather than one session's result.
Sub-question coverage is what that record is read for. When a buyer question keeps producing answers built on a comparison page or a third-party review you are absent from, answer coverage shows you which sources the answer rested on. Whether the fix is a passage you are missing or presence in a third-party source is a judgement a person makes on the snapshot. Sometimes the honest read is that the question is not worth chasing. Your Brand Core holds the claims you are entitled to make and the proof behind each, so a recommended action names a specific asset supported by evidence you already have.
Drafts are produced from approved claims and a person reviews everything before it publishes. After the asset goes live, the next analysis cycle reruns the same questions and compares the new snapshots against the old ones, which is how you learn whether the work was picked up or ignored.
What Trovance will not do is promise a citation in AI Mode or any control over what Google's fan-out retrieves. Nobody can offer that honestly, because the sub-queries are not exposed, the engines are probabilistic, and your competitors are publishing at the same time. There is also no single visibility score here, because presence and recommendation fail for different reasons, and averaging them hides the one you can act on.
What should you do this week?
Confirm the entry condition first, since nothing downstream matters without it. Check that your key pages are indexed and eligible to show a snippet, and that no inherited snippet configuration is quietly excluding the paragraphs you most want quoted. That takes an hour and rules out the failure mode no amount of writing fixes.
Then do the decomposition. Take your three highest-value buyer questions, write the sub-questions each implies, and mark which ones your site answers in a single self-contained paragraph. Fix those retrieval gaps before commissioning anything new.
Be honest about the timeline while you do it. Indexing and snippet fixes can surface within weeks, evidence improvements follow recrawling, and earning presence in third-party sources takes months, with measurement variance sitting on top of every reading. If you would rather watch the answers change than sample them by hand, start a free Trovance analysis and see which sub-questions your pages are already answering.
How the engines choose sources
How to get featured in Google AI Overviews - the surface this article's mechanics apply to first.
Google rankings versus AI citations - why the top 10 and the citation set have come apart.
What actually gets cited by AI engines - the page traits that correlate with being pulled in.
Where AI citations come from - how much of an answer sits off your own domain.
Make your passages retrievable
Do headings help AI retrieve long documents? - structure at the level retrieval reads.
Can likely buyer questions improve AI retrieval? - the decomposition exercise, tested.
AI crawlers don't run your JavaScript - the fetchability difference between Google and everyone else.
How to show up in ChatGPT without chasing hacks - the same evidence-first argument on another engine.
FAQs
How does Google AI Mode choose which sources to cite?
Google states that AI Mode and AI Overviews use a query fan-out technique, issuing multiple related searches across subtopics and data sources, then assembling an answer from what those searches return. A page must be indexed and eligible to show a snippet, and Google names no other requirement anywhere in its published guidance.
Do I need special schema or a new file to appear in AI Mode?
No. Google's documentation says there are no additional requirements and no new machine-readable files or markup needed for AI experiences. Structured data still helps Google parse a page for supported features, and the supported-types gallery contains no AI Mode entry. Indexing and snippet eligibility remain the two stated conditions.
Why do pages outside the top 10 get cited by AI Overviews?
Because fan-out searches sub-questions, not the question you typed. Ahrefs found 38% of AI Overview citations rank in the organic top 10, down from roughly 76% in July 2025, across 4 million AI Overview URLs. A page sitting outside the top 10 can still hold the passage that settles a hidden sub-search.
Can I stop Google AI Mode from using my content?
Partly, and only by subtracting. The single lever Google's documentation describes is snippet eligibility, so removing a page's eligibility removes it from the answers built on snippets. Nothing works in the other direction: no documented directive or markup requests inclusion in a Google AI experience.
Does rewriting my pages in a conversational style help?
The evidence says rarely. C-SEO Bench found only 3 of 54 tested unilateral conditions produced statistically significant citation-rank gains across two tasks and six domains. It re-tested the method families an earlier GEO study reported gains for and did not reproduce them, so treat rewrite tactics as unproven.
How many times should I run a query before trusting the result?
At least ten times, across several days, before treating any single answer as a reading. SparkToro's research, Search Engine Land's write-up of it, and a 2026 variance-components study all concern how much AI recommendations move between runs. Appearance rate across repeated runs is the honest measurement, and one session is a sample.
Should I optimize my own site or work on third-party sources?
Both, in that order. Profound found 57% of AI citations point to sources outside the brand's own site, so most of the citation set sits outside your control, though the share of citations is not the share of the answer. Fix indexing and passage-level answers on pages you own first.



