TL;DR
⏱️ The clock belongs to the engines: GPTBot's share of crawl requests moved from 2.2% to 7.7% in a year, so nobody can quote you a date for a citation.
🎯 The measurement is worth the effort because 51% of software buyers now begin research inside an AI chatbot.
🧪 Piece counts are folklore: C-SEO Bench found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains.
🔎 Retrieval gates the fast path: 87% of SearchGPT citations matched Bing's top results, so a page missing from conventional indexes is starting from behind.
📚 Check your baseline sources before you publish, because 57% of AI citations point to sources brands do not control.
No number of published pieces produces a citation on a schedule. The claim that twelve pieces, or fifty, flips an engine into citing you is folklore from an era when crawlers arrived on a predictable cadence. Time to citation is not a forecast you can buy. It is a measurement you take, and you can only take it if you wrote down the questions and captured the baseline before you published anything: define the question set, record how the engines answer it now, publish, rerun the identical questions, and compare the delta against the run-to-run spread you already measured.
The stakes justify that much rigor. G2's research found AI chatbots are now the single largest influence on B2B shortlists, and 51% of software buyers now begin research inside an AI chatbot. Verification is harder than it was in search because the thing you are measuring moves without you: SparkToro's research found AI engines are highly inconsistent when recommending brands, and Search Engine Land's write-up of it reports that recommendation lists rarely repeat exactly. One answer before and one answer after tells you nothing about your work.
Why can't anyone tell you how long it takes?
Because the clock belongs to each engine's crawl and retrieval pipeline, and those run on schedules you do not set. OpenAI documents four separate crawler user agents doing different jobs, and its help documentation describes ChatGPT search issuing one or more targeted queries at answer time, so what a buyer sees is assembled fresh on each ask. Cloudflare measured GPTBot's share of crawl requests rising from 2.2% to 7.7% across a year, and its bot report puts AI crawlers at 52% of crawler requests. Engine behavior sets the cadence, and that cadence keeps moving.
There are also two clocks running at once, and they have different speeds. The live retrieval path can surface a page quickly, and Seer found 87% of SearchGPT citations matched Bing's top results, so presence in a conventional index is strongly associated with the fast path, though the overlap is not complete. The slow clock is what the model already believes about your category. A 2026 analysis of brand dynamics in LLM recommendation systems examines how those prior associations form and shift; our reading of it is that they are sticky and spread unevenly across brands. A page can enter the fast path in days and still lose the recommendation for months.
Before either clock starts, there is a mechanical prerequisite that turns "how long" into "never". Vercel's crawler research with MERJ found the OpenAI and Anthropic crawlers fetch raw HTML without executing JavaScript, while Google's path uses the same Web Rendering Service capabilities and renders the page. Cloudflare's agent-readiness assessment of the 200,000 most visited domains scored how much of each site an agent can actually read. If your key pages render client-side, they are invisible to the crawlers that skip JavaScript, so any timing question about those engines is undefined until you fix it.
Does publishing a fixed number of pieces trigger citations?
No, and the benchmark evidence argues against the whole framing. C-SEO Bench (NeurIPS 2025) tested ten conversational-SEO rewrite methods and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains, across two tasks and six domains. If a rewrite tactic mostly fails to move citation rank once, repeating it twelve times mostly fails twelve times. Piece counts inherit whatever effect the underlying tactic has, and for most rewrite tactics that effect is indistinguishable from zero.
What does carry evidence is extractable proof on the page. The GEO study (Aggarwal et al., KDD 2024) found that adding statistics, quotations, and citations lifted a source's visibility inside generated answers by up to 40% on its benchmark, while keyword stuffing did nothing in the same error-bar table. That is visibility in the answer, not a citation rate, and it is a property of how a page is written.
Piece-count promises also fail a plainer test: they are stated in a form that cannot be checked. They never name the questions the citation would answer, the engines it would appear in, or how many runs would count as evidence. Ask for those three things and the promise either becomes a testable claim or evaporates. "We shipped twelve" and "we got cited" are two facts with no measured line between them.
What has to be defined before you publish?
The question set, written down and frozen, before anything ships. Five to fifteen questions phrased the way a buyer would actually type them, covering the moment the shortlist forms. 6sense found buyers work with roughly five vendors and already know 3.8 of them when the process starts, so the questions that matter are the ones that assemble that list.
Then define the engines and the run count, because each engine is its own measurement and cannot be averaged with the others. Semrush's AI Visibility Index analyzed 126 million AI search prompts to say anything general about engine behavior. You are measuring one brand against one frozen question set, so repetition is what protects you from noise.
Define the outcome you are recording, at the level of the rung. A mention is your name in an answer, a citation is your page used as a source, a recommendation is the engine advising the buyer to choose you, and a shortlist position is surviving the narrowing. The IAB's measurement work separates presence, prominence, portrayal, and persuasion for the same reason, and its guidance sorts brands into 4 groups by how they should approach the problem. Recording only "did my name appear" throws away the rung you are actually paid for.
How do you capture a baseline that survives noise?
Run every question at least ten times per engine, spread over several days, and record an appearance rate across those runs. A 2026 variance-components study documents how much of the variation between runs is noise. Ten runs per question per engine is our own working floor, and the paper prescribes no such number. Work on quantifying uncertainty in AI visibility makes the case for confidence intervals, and the same analysis reports medians that differ by platform, which is why a cross-engine average hides more than it shows.
Record the sources each answer used alongside the verdict. Profound's citation research found 57% of AI citations point to sources brands do not control, so a baseline built on review sites and community threads is telling you in advance that publishing on your own domain may not move it.
Freeze the wording, including capitalization and any product qualifier. Changing the question between baseline and re-run destroys the comparison, and small differences in how an entity is named change what gets retrieved: entity-oriented retrieval research across 443 configurations shows how much retrieval quality depends on resolving the entity cleanly.
How do you tell a real change from run-to-run variance?
Compare the delta to the spread you already measured, and refuse to interpret anything smaller. If your baseline appearance rate for a question was 2 of 10 and the re-run gives 3 of 10, one run of difference is well inside what a ten-run sample produces on its own. If it moves from 1 of 10 to 7 of 10 and holds across two more re-runs on different days, you have something worth acting on.
Keep a hold-out group. Pick two or three questions in the same category that you deliberately do not write for, and re-run them alongside the ones you targeted. If the hold-out questions move too, the market moved, or the engine did, and your content is not the explanation.
Re-run on a fixed interval, set in advance. Measurement systems choose an interval on purpose: Cloudflare Radar exposes intervals from 1-day requests down to 15-minute buckets so a comparison is always like for like. Weekly or biweekly is enough for most categories; irregular sampling turns every result into an anecdote.
What counts as movement, and what does not?
Movement is a change in appearance rate on the frozen question set, on a named engine, that exceeds your measured spread and persists across re-runs. A screenshot is not movement. Neither is one answer that happens to name you.
Watch the rung as well as the rate. Getting cited more while getting recommended the same amount is a real finding, and it usually means your evidence is being used to support an answer that still ends with someone else's name. That is a different problem from invisibility, and the fix is comparative proof.
Connect the loop to the business without pretending the attribution is clean. Referrer strings from AI answers are frequently empty, so treat any AI-referral count as a floor and lean on conversion behavior on the pages the questions point at: a key event rate is sessions with a key event divided by total sessions, and it tells you whether the arrivals are the right ones. Anyone offering a single number that collapses citation, recommendation, and revenue into one score is selling a simplification the measurement does not support.
How does Trovance run this verification loop?
Trovance holds the loop as a system rather than an afternoon of manual runs. You define the market questions your buyers actually ask, and it runs them repeatedly across AI engines, preserving each answer run with its full context: who was mentioned, who was cited, who was recommended, and which sources carried the answer. Because the questions stay frozen and the runs keep accumulating, the baseline is not a document someone has to remember to write.
The preserved answer snapshots are what make a delta legible. Answer coverage across your question set is a rate with a history behind it, so when it moves you can look at whether the sources changed or whether a competitor's page entered the answer. Recommended actions name the specific asset the record says is missing, and your Brand Core holds the claims you are entitled to make with the proof behind each one, so drafts are produced from approved claims and a person reviews everything before it publishes.
The analysis cycle is the part that answers "how long". After an asset goes live, the next cycle reruns the same questions and compares the result to everything recorded before it, which turns time to citation into an observed interval for your brand on your own questions. Over successive cycles you learn your own variance, which is the number that decides whether any given change was real.
What Trovance will not promise is a date or a guaranteed citation. The engines are probabilistic and the sources shift underneath you, so any system that names a timeline is asserting control it does not have. The platform preserves enough evidence that you can tell the difference between a change and a coincidence, and it is building toward tighter attribution between answer movement and downstream behavior.
What should you do this week?
Write the question set first, before any content decision. Ten questions, phrased as a buyer would ask them, covering the shortlist moment. Then check the mechanical prerequisite: fetch your key pages with JavaScript disabled and confirm an engine reading raw HTML would find an answer there. If it would not, fix that first; every timing question downstream is meaningless until you do.
Next, run the baseline before you publish. Ten runs per question per engine, across at least three days, recording the rung and every source cited. Add two hold-out questions you will not write for. Only then ship the work, and set a fixed re-run date so the comparison exists on a calendar rather than in someone's memory.
Report the ordering and let your own re-runs supply the durations. Retrieval fixes are the fastest to show up because they act on the live path, evidence improvements wait on re-crawling, and presence in third-party sources is slowest because you do not control the publisher. Your re-runs are what turn that ordering into numbers. If you want the loop running continuously without rebuilding it by hand each quarter, start a free Trovance analysis and let the baseline accumulate while you work.
Build the measurement
How to track AI citations · the record-keeping that makes a delta legible.
How to measure AI search visibility without one score · what to measure when no single index is honest.
There is no such thing as an AI visibility score · why that number cannot exist.
AI share of voice · a rate worth tracking, and when it misleads.
Fix what the measurement finds
AI crawlers don't run your JavaScript · the prerequisite behind every timing question.
Where AI citations come from · which sources carry the answers you want to move.
The GEO playbook for getting cited by AI engines · tactics worth testing once you can measure them.
A visibility gap is not a content brief · why a measured gap is not automatically an article.
FAQs
How long does it take for AI engines to cite a new page?
There is no fixed interval. Timing depends on each engine's crawl and retrieval cadence, which you do not control. OpenAI documents four separate crawler user agents doing different jobs. The only honest answer comes from rerunning a frozen question set after publishing and comparing appearance rates to your recorded baseline.
How do I know my GEO work is working?
Compare appearance rates on the same questions before and after publishing. One answer on each side proves nothing. Run each question at least ten times across several days at baseline, publish, then rerun the identical wording. A change smaller than the run-to-run spread you measured at baseline is noise you should refuse to interpret.
Does publishing twelve pieces get you cited?
No study supports a piece count as a trigger. C-SEO Bench tested ten conversational-SEO methods and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. Repeating a tactic with little measured effect does not accumulate into one. Extractable evidence on the pages that answer buyer questions has better support.
How many times should I run each question?
Ten runs per question per engine is our working floor, spread across several days. A 2026 variance-components study documents how much of the run-to-run variation is noise, though the ten-run figure is ours and the paper prescribes no count. Fewer runs produce a figure that looks precise and cannot support the comparison you want.
Can one AI visibility score tell me whether GEO is working?
No. Mentions, citations, recommendations, and shortlist survival fail for different reasons and move on different timelines. The IAB's measurement work separates presence, prominence, portrayal, and persuasion for that reason. Collapsing them into a single index hides which rung moved and leaves you unable to choose the next action.
Why did my answer change without me publishing anything?
Because the engines change on their own. SparkToro's research found AI engines are highly inconsistent when recommending brands, with the same prompt returning different vendor lists across runs. Competitors publish, third-party sources update, and models are retrained. This is exactly why a baseline needs many runs before you attribute anything to your work.
What should I record in a baseline?
The exact question wording, the engine, the date, whether you were mentioned, whether your page was cited, whether you were recommended, and every source the answer used. Profound found 57% of AI citations point to sources brands do not control, so the source list often explains the result better than your own pages.



