TL;DR
📄 Google names 3 things it does not require for AI Overviews or AI Mode: additional requirements, special optimizations, and new machine readable files or markup.
🔍 No published source measures llms.txt against citation outcomes, and a 2026 variance-components study found run-to-run noise large enough to swamp real differences in small samples.
🧱 AI crawlers do not execute JavaScript, and Cloudflare measured agent readiness across the 200,000 most visited domains: a page whose copy is missing from the raw HTML is unreadable to those crawlers no matter what file sits at your root.
🚦 Access is documented and decidable: OpenAI publishes four relevant user agents and pay-per-crawl answers a crawler with HTTP 402.
🧪 C-SEO Bench found only 3 of 54 unilateral conditions produced significant citation-rank gains, so budget an hour on llms.txt as a cheap bet.
No, your site does not need llms.txt to appear in Google's AI Overviews or AI Mode, and that is not a matter of opinion. Google's documentation for its AI features states there are no additional requirements to appear, no special optimizations, and no new machine readable files, AI text files, or markup. The condition it does state is ordinary: the page has to be indexed and eligible for snippets. If you were hoping llms.txt was the missing switch on Google's surfaces, the vendor that operates those surfaces has already answered the question in writing.
That answer covers Google and only Google. For every other engine, the honest position is that you cannot verify adoption from outside, and no published source measures what the file does to citation rates. The useful question is what an hour spent on it costs you elsewhere. Google's separate guidance on succeeding in AI search points at the same ordinary mechanics, its helpful content guidance was not amended to accommodate a new file, and its AI Mode announcements describe a query fan-out over the existing index rather than a new intake channel.
What is llms.txt proposed to do?
llms.txt is a proposed convention: a markdown file at the root of your domain listing your important pages with a short description of each, so a model can read a curated map instead of crawling the whole site. The appeal is real. Retrieval is expensive and context windows are finite, so a short index is cheaper for a machine to read than a sitemap plus two hundred templated pages.
The shape is familiar. robots.txt tells crawlers where not to go, a sitemap tells them what exists, and llms.txt would tell them what matters and why. Nothing about the design is unserious, and the people advocating for it are not confused about how retrieval works.
The problem sits one layer down. A convention only does something when the software on the other end reads it, and reading is the part you cannot observe.
Compare it with markup that engines document. Google publishes what structured data is used for and names the specific types Search supports, while the vocabulary itself is published by schema.org. llms.txt has advocacy. It does not yet have documentation like that.
Can you verify that any engine actually reads it?
Not from your own logs, and not reliably. You can see a request for /llms.txt arrive, which proves something fetched the file. It does not prove the file changed an answer, and no source in the public record measures the effect of publishing one on how often you get cited. Anyone quoting you a citation lift from llms.txt is quoting a number nobody has produced.
The measurement problem runs deeper than a missing study. AI answers are unstable on their own terms. SparkToro found AI engines are highly inconsistent when recommending brands, a separate study found recommendation lists rarely repeat exactly, and a 2026 variance-components study found run-to-run noise large enough to swamp real differences in small samples.
Ship a file on Monday, count more mentions on Friday, and you have learned nothing about the file. Work on quantifying uncertainty in AI visibility puts confidence intervals around the same problem, and the intervals are wide. Treat "llms.txt lifted our citations" and "llms.txt did nothing" as equally unsupported claims, because both of them are.
What is documented, and therefore worth your hour instead?
Three things about machine access are published by the vendors themselves, which puts them in a different category from a convention you have to take on faith.
Crawler identity, which robots.txt already controls
Google publishes its crawler and user agent list, and OpenAI documents four relevant user agents with distinct jobs. That distinction is the practical one: a training crawl and a live search fetch are different transactions, and you can allow one while refusing the other. Both vendors read robots.txt, and both say so.
Whether your content is in the raw HTML at all
Vercel's crawler research with MERJ found the AI crawlers do not execute JavaScript, so client-rendered copy reaches them as an empty shell. That is the fact this question turns on. A curated map cannot help a page that returns an empty container, and a perfect llms.txt on a site like that is a map to nothing.
Access, which is now a commercial decision
Cloudflare's pay-per-crawl uses HTTP 402 as an answer to a crawler, and its measurement of AI crawler share rising from 2.2% to 7.7% shows why the question arrived when it did. A later bot report puts 52% of crawler requests in the AI category. Whether you allow, charge, or refuse decides far more about machine access than any file enumerating your pages.
Why does the tactic reflex mislead people here?
llms.txt belongs to a family of moves that feel like optimization and mostly are not. C-SEO Bench tested conversational SEO methods across two tasks and six domains and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. The NeurIPS 2025 version of that work is public, and the code and per-condition results are on GitHub. The rewrite tricks that dominate advice threads mostly did not survive a controlled test, and a text file at your domain root has less evidence behind it than the tricks that failed.
What the record does support is duller. 38% of AI Overview citations rank in the organic top 10 across 4 million AI Overview URLs in March 2026, down from roughly 76% in July 2025. 87% of SearchGPT citations across a 500-citation sample matched Bing's top results, and 57% of AI citations point to sources brands do not control, which is an argument for presence in the third-party sources engines cite. Retrievability through the ordinary index is the part with published measurement behind it, and those numbers show it is decreasingly sufficient on its own.
The page-level correlates are documented too. An analysis of 548,534 pages mapped which traits track with being pulled into an answer, and ChatGPT search runs one or more targeted queries against a live index. The web search tool documentation notes results can be cached for around 10 minutes, which is another reason a Friday reading proves little.
So should you publish one anyway?
Publish it if you want to, and budget it as a cheap bet. The file takes an hour and breaks nothing. If a future agent stack does read it, you are already there. That is the whole case, and it is a reasonable one, but it includes no expectation that citations move.
If you ship one, write it for a reader with no patience. Use the exact terms a buyer would use, and list only URLs that resolve. A stale llms.txt is worse than none, because it hands an agent an outdated map of a site that has since changed, and you will not get a bounce message telling you so. Refresh it whenever your page structure shifts.
The failure mode to avoid is substitution. An hour on a text file is fine. A quarter spent debating machine-readable conventions while your pricing page renders client side, your robots.txt blocks a fetcher you meant to allow, and your claims sit on pages with no evidence attached is a quarter traded for the least documented item on the list. The IAB's framework for measuring visibility in the AI era organizes the work into 4 groups for brands, and none of them is a file.
How does Trovance treat questions like this one?
Trovance is built for the part you can actually observe. You define the questions your buyers ask, and the platform runs them repeatedly across AI engines, preserving each answer run as a snapshot with its full context: who was mentioned, who was cited, who was recommended, and which sources carried the answer. Answer coverage across those tracked questions is a measurement.
That record is what makes a change diagnosable. When you publish an asset, fix a rendering problem, or adjust access rules, the next analysis cycle reruns the same questions and compares the new snapshots against the stored ones, so you can see whether anything moved and whether the movement exceeds normal run-to-run variance. Your Brand Core holds the claims you are entitled to make and the proof behind each, and recommended actions name the specific asset the evidence record says is missing.
Trovance will not promise that llms.txt changes an answer, because nothing in the public record supports that either way, and the platform does not produce a single universal visibility score that would let a file appear to move a number. It cannot control what a model says, and it does not publish without a person approving the draft. What it does is preserve enough evidence that you can tell a real shift from noise.
The boundary is worth stating plainly. If your pages are unreadable to a crawler or your claims have no proof behind them, no measurement layer repairs that, and Trovance will show you the gap instead of papering over it. The value is in the honest before-and-after.
What should you do this week?
Sequence it by how much documentation stands behind each step. First, fetch your five most commercially important pages with JavaScript disabled and read what comes back; if the copy is missing, stop and fix that. Second, open your robots.txt and check it against the published user agent lists, then decide deliberately which fetchers you allow. Third, put extractable proof on the pages that answer buyer questions; an analysis of 548,534 pages mapped which traits track with being pulled into an answer.
Fourth, and only fourth, write your llms.txt if you want one. Give it an hour, keep it current, and do not report it to anyone as a visibility initiative. Then measure the way the variance research demands: repeated runs of the same buyer questions over time.
If you want that measurement running continuously instead of by hand, start a free Trovance analysis and see which of your buyer questions already have an answer without you in it.
Make your site readable first
What an agentic browser sees on your site - the view that decides whether any of this matters.
AI crawlers don't run your JavaScript - the fetchability problem underneath every convention.
Technical SEO for AI search - the documented mechanics, in order.
Schema markup and AI citations - what markup engines actually document using.
Measure it honestly
How to measure AI agent traffic - separating a crawl from an audience in your own logs.
How to show up in ChatGPT without chasing hacks - the tactics that survived testing.
There is no such thing as an AI visibility score - why a composite number hides the real movement.
Measuring AI search visibility without one score - the measurement design this article assumes.
FAQs
What is llms.txt and does my site need it?
llms.txt is a proposed markdown file listing your key pages with short descriptions for models to read. Google's AI features documentation states no new machine readable files or markup are needed to appear in AI Overviews or AI Mode. Publish one if you like, but treat it as optional housekeeping.
Does Google read llms.txt?
Google's documentation says appearing in its AI features carries no additional requirements, no special optimizations, and specifically no new machine readable files, AI text files, or markup. The stated conditions are that a page is indexed and eligible for snippets. That is a direct answer for Google surfaces and nobody else's.
Will publishing llms.txt increase my AI citations?
Nobody can say. No published study measures citation outcomes from publishing the file, and AI answers vary enough between identical runs that a before-and-after reading on one week proves nothing. Claims in either direction are unsupported. Budget the hour as a cheap bet and expect nothing measurable from it.
What should I do instead of writing llms.txt?
Check that your commercially important pages contain their copy in the raw HTML, since research found AI crawlers do not execute JavaScript. Then review robots.txt against the published crawler user agent lists from Google and OpenAI. Both are documented by the vendors and both gate everything downstream.
Is llms.txt the same as robots.txt or a sitemap?
No. robots.txt and sitemaps are long-standing conventions that major crawlers document reading, and access decisions made there take effect. llms.txt is a newer proposal for a curated site summary, and no engine documents using it. The shapes are similar; the standing behind them is not.
Do rewrite tactics for AI search work better than llms.txt?
Neither has evidence in its favor. C-SEO Bench tested conversational SEO methods across two tasks and six domains and found only 3 of 54 unilateral conditions produced statistically significant citation-rank gains. No published study measures llms.txt at all, so it cannot be ranked against them. Retrievability through the ordinary index is the part with published support.
How often should I update llms.txt if I publish one?
Whenever your page structure or key URLs change. A stale file gives an agent an outdated map of a site that has moved on, and no error message tells you it happened. Keep descriptions literal, list only URLs that resolve, and review it alongside your sitemap.



