AI Visibility

llms.txt: does it actually work? What the server logs and 300,000-domain studies say

No engine ever committed to reading it. 97% of files never get fetched. The real consumers are coding agents. Here is the honest verdict and what actually moves citations.

Editorial line illustration of a single plain text file at a crossroads of crawler paths, most paths bypassing it, one narrow path reaching it labelled coding agent, with a single orange accent shape

► Bottom line up front

Mostly no, and that is fine. No AI engine has ever committed to reading llms.txt. In Ahrefs' June 2026 study of 137,210 domains, 97 percent of llms.txt files received zero requests of any kind, and SE Ranking's analysis of roughly 300,000 domains found no effect on AI citation frequency. The file costs nothing to ship correctly at build time, its heaviest real consumer is coding agents rather than answer engines, and the levers that actually move citations are schema and extractable content. Ship the file, then spend your effort where the evidence is.

What the llms.txt specification actually requires

llms.txt is a single Markdown file at your domain root that offers a large language model a curated map of your site, and its specification asks for far less than most people assume. Jeremy Howard of Answer.AI proposed it on 3 September 2024, and the canonical specification lives at llmstxt.org. The whole thing is a good deal smaller than the volume of argument that has grown up around it over the past two years.

The specification defines exactly one required element: an H1 heading with the name of the site or project. That is it. As the spec puts it, that H1 "is the only required section." Everything else is optional: a blockquote summary of what the site is, then H2-delimited sections listing links with short descriptions. You could satisfy the specification with a single line of Markdown. That minimalism is deliberate, and it is also why the file is cheap to ship and easy to get wrong.

llms.txt is the map, llms-full.txt is the territory

The two files get conflated constantly, and only one of them is in the specification. Only one. llms.txt is the lean index: H1, optional summary, optional link sections. llms-full.txt is not in the specification at all, which is the part that trips people up when they cite it as though the standard blessed it, because the standard did no such thing and instead reaches for a different mechanism entirely. It is an ecosystem convention, popularised by documentation tooling, that concatenates a site's entire content into one Markdown bundle for one-shot ingestion. The specification itself takes a different route to full content: per-page Markdown mirrors, one clean .md file alongside each HTML page. Three surfaces, then. So when someone says a site "has llms.txt," ask which of these three they actually mean.

The famous adopter list is mostly a tooling default

The strongest-looking argument for llms.txt is the roster of serious companies that ship one. Anthropic, Cursor, Stripe, Vercel, and Hugging Face all publish llms.txt files. Impressive list. That looks like a considered industry bet, right up until you trace how the files actually got there and find that a single documentation platform's default setting explains most of the names on it. Mintlify, the documentation platform many of them run, rolled out automatic llms.txt generation on 14 November 2024, per mintlify.com. Much of the celebrated adoption is that default firing, not a deliberate visibility decision. A cascade of auto-generated files is weaker evidence than a list of names implies. The roster is a tooling default, not a verdict.

Does any AI engine actually read llms.txt?

No major AI engine has committed to reading third-party llms.txt files, and the ones that publish their own do so as a byproduct rather than a policy. Google has said so plainly. OpenAI, Anthropic, and Perplexity have said nothing that amounts to a commitment. The one demonstrable consumer in server logs is a coding agent, not an answer engine.

AI engine commitment to reading llms.txt, August 2026WHO COMMITS TO READING llms.txt • AUGUST 2026GoogleNOMueller: "comparableto the keywordsmeta tag."Illyes, Jul 2025:not supported,not planned.OpenAI + AnthropicNO COMMITMENTBoth publish theirown llms.txt for docs.GPTBot: 4.51% ofllms.txt requests.ClaudeBot: 0.80%.Incidental, not policy.PerplexityABSENTNo official statement.PerplexityBot absentfrom llms.txt requestlogs entirely.Coding agentsREAL CONSUMERClaude Code out-fetched every AIretrieval andtraining crawler.Developer tools, notanswer engines.Sources: Mueller (Reddit, Apr 2025) · Illyes (Search Central Live, Jul 2025) · Ahrefs log study, Jun 2026.
Engine-by-engine commitment, August 2026. Google says no on the record; OpenAI, Anthropic, and Perplexity have made no commitment; the only demonstrable consumer in server logs is a coding agent. Crawler shares are of requests to the small fraction of files that got any traffic.
Engine-by-engine commitment to reading third-party llms.txt, August 2026
EngineOfficial statementCrawler behaviour in logsVerdict
GoogleMueller (Apr 2025): "comparable to the keywords meta tag." Illyes (Jul 2025): not supported, not planned.Googlebot absent from llms.txt request logs.No
OpenAINo statement that GPTBot or OAI-SearchBot reads third-party files. Publishes its own for docs.GPTBot 4.51% of llms.txt requests. Incidental.No commitment
AnthropicNo statement. Own file exists via the Mintlify default, not a crawler policy.ClaudeBot 0.80% of llms.txt requests.No commitment
PerplexityNo official statement. Users can paste a file URL as context manually.PerplexityBot absent from log breakdowns.No
Coding agentsNot an answer engine. Fetches files to build site context.Claude Code out-fetched every retrieval and training crawler.The real consumer

Google: on the record, twice, that it does not use it

Google has been the clearest. John Mueller wrote on Reddit in April 2025 that, as far as he knew, none of the AI services had said they were using llms.txt, and added: "To me, it's comparable to the keywords meta tag," a reference to a signal Google abandoned as spam-prone. Gary Illyes reinforced it at Search Central Live in July 2025, saying Google does not support llms.txt and is not planning to. Googlebot is absent from llms.txt request logs entirely.

The December 2025 Google confusion, and its retraction

There was one wobble worth naming so you can discount it. In December 2025 some Google-owned properties were briefly found hosting llms.txt files, and the SEO community read that as a quiet endorsement. It was not. Mueller clarified publicly that the files were not a signal of Google using llms.txt for ranking or citation, and the files were removed. Google's position through August 2026 is unchanged: it does not read your llms.txt.

OpenAI, Anthropic, and Perplexity: silence, byproducts, and absence

The others resolve quickly. OpenAI has made no statement that GPTBot or OAI-SearchBot reads third-party llms.txt, and GPTBot's 4.51 percent share of llms.txt requests reads as incidental crawling. Anthropic publishes its own llms.txt for its docs at platform.claude.com, but that came through the Mintlify default, not a crawler policy, and ClaudeBot is 0.80 percent of llms.txt requests. Perplexity has said nothing, and PerplexityBot is absent from the log breakdowns. So far, so bleak. Then one name stands out. The one entity that genuinely fetches these files is Claude Code, the coding agent, which out-fetched every retrieval and training crawler in Ahrefs' logs. The real audience for llms.txt is developer tools, not the answer engines you are trying to get cited in.

The numbers: 97 percent of llms.txt files never get fetched

Two studies, two methods, same landing spot: llms.txt is barely fetched and does not measurably move citations. The Ahrefs log study measures whether the file gets read at all. The SE Ranking modelling study measures whether it does anything for the small fraction of sites where it is.

Who actually fetches llms.txt, Ahrefs June 2026WHO FETCHES llms.txt • AHREFS, 137,210 DOMAINS, JUN 2026The headline: 97% of files got zero requests97% of llms.txt files: ZERO requests of any kindOf the 3% that got ANY request, share by fetcher:SEO audit tools21.7%AI retrieval bots19.5%GPTBot (OpenAI)4.51%ClaudeBot0.80%DeepseekBot0.02%Read this: SEO tools, not answer engines, lead the fetches. AI retrieval is second, on 3% of files.Source: Ahrefs, Louise Linehan and Xibeijia Guan, June 2026 (May 2026 logs).
Ahrefs log study, 137,210 domains, June 2026. The dominant fact is the red bar: 97 percent of files were never fetched. Among the 3 percent that were, SEO audit tools led at 21.7 percent, ahead of AI retrieval bots at 19.5 percent.

The Ahrefs study: 97 percent of files never get fetched

Louise Linehan and Xibeijia Guan of Ahrefs analysed 137,210 domains against May 2026 server logs, published in June 2026, per ahrefs.com. The headline finding is blunt: 97 percent of llms.txt files received zero requests of any kind. Not zero AI requests. Zero requests, full stop. Of the 3 percent that saw any traffic at all, the largest single fetcher was not an answer engine and not a training crawler but the class of SEO audit tools running site scans, which took 21.7 percent of those requests, ahead of the AI retrieval bots on 19.5 percent. GPTBot sat at 4.51 percent, ClaudeBot at 0.80 percent, and DeepseekBot at 0.02 percent. A file that almost nobody requests cannot be doing much work.

The SE Ranking study: no measurable citation effect

The second study asked the harder question. Not who fetches the file, but whether it changes the outcome for the sites that have one, which SE Ranking tested by analysing roughly 300,000 domains, published on 20 November 2025 via Search Engine Journal, and controlling throughout for authority, schema, and recency so the file's contribution could be isolated from everything that usually travels with it. Their machine-learning model's accuracy actually improved when the llms.txt feature was removed. Noise, in other words. Their conclusion: llms.txt "doesn't seem to directly impact AI citation frequency. At least not yet." If a controlled model gets better at predicting citations by ignoring your llms.txt, the file is not your lever.

What the same data says does move citations

The other side of the SE Ranking study is where the effort points. What correlated, and by how much: In the same 300,000-domain dataset, FAQPage schema correlated with roughly 34 percent more Perplexity citations and 28 percent more in ChatGPT, and ClaimReview markup on statistic-dense content correlated with about 41 percent more citations in Google AI Mode, which together tell you that answer engines reward pages that pre-structure their claims into extractable question-answer and fact-review units the machine can lift without guessing. Organization sameAs added around 22 percent. Speakable added around 18 percent in AI Mode. These are structured-data and extractable-content levers, which is to say they are the machine-readable substance of the page rather than a courtesy note filed at the root, and they are where every scrap of measured signal in that dataset actually lived.

Citation levers by measured correlation, SE Ranking 300K-domain study
LeverMeasured correlation with citationsSource
FAQPage schemaApprox. +34% Perplexity, +28% ChatGPTSE Ranking, Nov 2025
ClaimReview on stat-dense contentApprox. +41% Google AI ModeSE Ranking, Nov 2025
Organization sameAsApprox. +22%SE Ranking, Nov 2025
Speakable specificationApprox. +18% AI ModeSE Ranking, Nov 2025
llms.txt file presentNo measurable effect; accuracy improved when removedSE Ranking, Nov 2025

Our experience matches the studies. Our llms.txt has never been our citation bet. We build it, we ship it, and the effort we actually invest goes into schema and clean extractable copy, the levers the 300,000-domain data says move AI mentions. The map file is hygiene. If you want the mechanism behind those levers, the Generative Engine Optimisation guide walks the schema architecture, and the brand citation measurement stack covers how to attribute a citation to a cause rather than a guess.

Why does Leapbuzz still ship llms.txt?

We ship it because the marginal cost is zero and the downside is nothing, not because we think it earns citations. Our build regenerates the whole stack on every deploy, so shipping a correct file is a settled decision rather than a recurring choice. The trick is to ship the format correctly, because a lazy stub is worse than useless.

The build already does it, so the decision is made

In practice: our scripts/build.py regenerates three surfaces on every single deploy: /llms.txt, /llms-full.txt, and a per-page /index.md Markdown mirror alongside every canonical page. The /llms.txt file follows the AnswerDotAI format properly: an H1 with the brand name, a blockquote that summarises what Leapbuzz is, then H2 link sections for home, services, platforms, industries, and blog. None of that is manual labour, and that distinction is the whole argument, because the moment a correct file requires a human to hand-maintain it the economics flip and the effort stops being worth the negligible return the studies describe. Once the generator is written, the cost of keeping a correct file live is a rounding error. That is the entire economic case. Zero effort, zero downside, small upside for the coding agents that do read it.

Ship the correct format, not a linkless stub

One failure mode, and it is common: shipping a file that technically parses but says nothing. The specification requires only the H1, so a file that is nothing but an H1 with no summary and no links is technically valid and practically a stub, the sort of box-tick that satisfies a validator while telling a reader precisely nothing about what the site is or where its useful pages sit. Lighthouse 13.4 flags present-but-linkless llms.txt files, so the minimal-compliant version is also the version a tool will mark against you. Ship the full correct shape: H1, blockquote summary, H2 link sections. It takes no longer to generate than the stub, and it is the version that is actually legible to whatever does read it.

Should you ship llms.txt? A decision flowSHOULD YOU SHIP llms.txt? • DECISION FLOWDoes your build generate it free?YESNOShip the CORRECT format:H1 + blockquote + H2 links.Not a linkless stub.Do not hand-build it.No effort beyond automation.The evidence does not justify it.Spend the real effort here instead:1. Schema (FAQPage, ClaimReview, sameAs, Speakable)2. Extractable chunks · 3. Cloudflare Content Signals prep
The decision is short. If the build generates llms.txt for free, ship the correct format and stop. Put real effort into schema, extractable content, and Content Signals preparation, the levers the studies actually measured.

The real 2026 standards fight is in robots.txt

The real 2026 fight over AI content access is inside robots.txt, the file every crawler already fetches and already obeys. Two separate efforts are converging there, and both will matter to your content long after the llms.txt debate has faded. Cloudflare Content Signals go live for every Cloudflare user on 15 September 2026, one month after this post. The IETF is drafting a competing vocabulary. Same base file. A marketing lead should be preparing for the first of those now.

Cloudflare Content Signals: live 15 September 2026

Cloudflare announced Content Signals in March 2026 and set them effective 15 September 2026 for all Cloudflare users, per TechCrunch. The mechanism is a Content-Signal field inside robots.txt with three categories: search, ai-input, and ai-train. A publisher can state that content may be used for search but not for AI training, for example. This matters more than llms.txt for one structural reason: it rides the robots.txt file that crawlers already fetch and already honour, rather than a separate file that 97 percent of the time nobody requests.

IETF AIPREF: drafts only, so far

The IETF is on the same track. The AIPREF working group, launched in January 2025, is drafting a standard vocabulary for AI-usage preferences; it produced working drafts through 2025, per Search Engine Land. As of August 2026, nothing is finalised. Drafts only. What matters for planning is the direction, and the direction is unambiguous once you line the two efforts up side by side: AIPREF, exactly like Cloudflare Content Signals, is built on robots.txt rather than a new file at the root. Both bet against the separate-file model that llms.txt represents. The industry is quietly consolidating AI access control back into the one file crawlers have honoured for thirty years.

What a marketing lead should actually do before 15 September

Turn the abstract into a checklist with a deadline attached. Three moves, in order.

  1. Decide your Content Signals policy before 15 September 2026. Work out, per content type, whether you want to permit search, ai-input, and ai-train. A publisher protecting premium content will differ from a brand chasing citation reach. This is a business decision, not a technical one, and it needs an owner.
  2. Keep shipping a correct llms.txt if your build already does. It costs nothing and serves the coding agents that read it. Do not add manual effort, and do not expect citations from it.
  3. Move your real budget to schema and extractable content. FAQPage, ClaimReview, sameAs, and Speakable are the measured levers from the SE Ranking data. That is where the citation work belongs. The visibility optimisation service is how we run this end to end across the five markets.

llms.txt is cheap insurance and a courtesy to coding agents, not a citation lever. No engine committed to reading it, 97 percent of files never get fetched, and a 300,000-domain model got better at predicting citations by ignoring it. Ship it if your build makes it free, then put your effort where it is measured. At Leapbuzz we run exactly that split across Singapore, Malaysia, Australia, the US, and Canada: correct files at build time, real work on schema and extractable content, and Content Signals policy decided ahead of the September deadline. That is the version of this advice I would stake a client engagement on.

Frequently asked questions

Does llms.txt actually work?

Mostly no. No AI engine has committed to reading third-party llms.txt files, and an SE Ranking analysis of roughly 300,000 domains published on 20 November 2025 found the file does not directly affect AI citation frequency. An Ahrefs study of 137,210 domains, published in June 2026, found 97 percent of llms.txt files received zero requests of any kind. It costs nothing to ship correctly at build time, so ship it, but do not expect citations from it.

Does ChatGPT use llms.txt?

OpenAI has made no official statement that GPTBot or OAI-SearchBot reads third-party llms.txt files. In Ahrefs' June 2026 log study, GPTBot accounted for 4.51 percent of requests to the small share of files that got any traffic, which reads as incidental crawling rather than declared policy. OpenAI does publish its own llms.txt for its developer documentation. So it ships one, but has not said its crawlers honour yours.

Does Google use llms.txt?

No. John Mueller of Google said in April 2025 that no AI service had confirmed using llms.txt and compared it to the keywords meta tag. Gary Illyes said at Search Central Live in July 2025 that Google does not support it and is not planning to. In December 2025 some Google-owned properties briefly hosted llms.txt files, causing confusion, and Mueller clarified it was not an endorsement. The files were removed.

How does llms.txt work?

llms.txt is a single Markdown file at your domain root, proposed by Jeremy Howard of Answer.AI on 3 September 2024. The specification at llmstxt.org requires only one thing: an H1 with the site or project name. Optional parts follow: a blockquote summary and H2-delimited sections of links. The idea is to give a large language model a clean, curated map of your site. Whether any engine reads it is a separate question, and mostly they do not.

What is the difference between llms.txt and llms-full.txt?

llms.txt is the lean index defined by the specification: an H1 plus optional link sections. llms-full.txt is not part of the specification at all. It is an ecosystem convention, popularised by tooling such as Mintlify, that concatenates a site's full content into one Markdown bundle for one-shot ingestion. The official specification instead defines per-page Markdown mirrors. So llms.txt is the map; llms-full.txt is the whole territory in one file.

Should I still create an llms.txt file?

Yes, if your build generates it for free. The downside is zero and the file is genuinely useful to the one confirmed consumer: coding agents. In Ahrefs' June 2026 logs, Claude Code out-fetched every AI retrieval and training crawler. Ship the correct format, an H1 plus a blockquote plus H2 link sections, not a linkless stub, which Lighthouse 13.4 flags. Then spend your real effort on schema and extractable content, which do move citations.

Is llms.txt the same as robots.txt?

No. robots.txt tells crawlers what they may access and is honoured by search engines. llms.txt is a proposed separate file that offers large language models a curated content map, and no engine has committed to reading it. The two are now converging in an important way: Cloudflare Content Signals, effective 15 September 2026, adds AI-preference fields inside robots.txt, and the IETF AIPREF working group is drafting a robots.txt-based vocabulary. The action is moving back into robots.txt.

What are Cloudflare Content Signals?

Cloudflare Content Signals are AI-usage preferences expressed inside robots.txt using a Content-Signal field with three categories: search, ai-input, and ai-train. Cloudflare announced them in March 2026 and they become effective on 15 September 2026 for all Cloudflare users. They let a publisher state, for example, that content may be used for search but not for AI training. Unlike llms.txt, this rides the existing robots.txt file that crawlers already read.

What actually improves AI citations?

Schema markup and extractable content, per SE Ranking's 300,000-domain analysis published 20 November 2025. In that data, FAQPage schema correlated with roughly 34 percent more Perplexity citations and 28 percent more in ChatGPT. ClaimReview on statistic-dense content correlated with about 41 percent more in Google AI Mode. Organization sameAs added around 22 percent and Speakable around 18 percent. llms.txt showed no measurable citation effect in the same model.

Why do Anthropic, Stripe, and Vercel all have llms.txt if it does nothing?

Because most of them did not decide to. Mintlify, the documentation platform many of them use, rolled out automatic llms.txt generation on 14 November 2024. Anthropic, Cursor, Stripe, Vercel, and Hugging Face all ship one, and much of that adoption is the Mintlify default rather than a deliberate visibility bet. The famous adopter list is largely a tooling cascade, which is why its presence is weaker evidence of effectiveness than it looks.

Related

Chasing AI citations, not just files?

We will show you which levers move your citation share and which are theatre.

Five-market capability across Singapore, Malaysia, Australia, the US, and Canada. 20-minute diagnostic call. Findings yours regardless.

Talk to Leapbuzz →