Published figures for llms.txt adoption run from "less than 0.005% of all websites" to 57%, and most of them are right about what they measured. They disagree because each one counts a different set of sites with a different test of what a file is. Crawls of ranked lists of large sites, made between May and September 2026 with a check that the response is a real file, found /llms.txt on 5.9% to 9.3% of them.
The most common way to get a higher number is to count every HTTP 200. Many sites answer any path with their home page or an error page and a 200 status. SEOmator's scan of the top 100,000 domains found 20,171 answering /llms.txt with a 200, and 8,598 serving a file: "a 2.35× overcount if you stop at the status line."
We read every public dataset we could find at its source on 29 September 2026. Below they are lined up by method, followed by the counts that do not measure adoption at all, and our own directory's figures, which update live.
llms.txt adoption: every public study, side by side
The llms.txt proposal is a convention, not a standard: Jeremy Howard published it on 3 September 2024 and revised it as version 2 in August 2026. It makes an H1 "the only required section" and lets the file live at the root "or at any subpath". Nobody governs how adoption is counted, so every study defines it for itself.
| Study | Data from | Sites measured | What counted as a file | Result |
|---|---|---|---|---|
| Rankability | Sep 2026 | Tranco top 1,000 and top 10,000 | 200, Markdown structure, and a control request to a made-up path to catch sites that answer everything; unreachable sites stay in the denominator | 9.3% of the top 1,000; 8.3% of the top 10,000 |
| SEOmator | Sep 2026 | Cloudflare Radar top 100,000 | 200, not HTML, JSON, a bot challenge or a robots.txt, at least 20 bytes | 8.6% of all; 11.4% of the 75,353 that answered |
| Chris Humphrey | Jun 2026 | Majestic top 10,000 | 200, "genuinely Markdown rather than HTML" | 7.4% |
| Thunderbit | May 2026 | Tranco top 10,000 | 200, final path still /llms.txt, not HTML, not empty | 5.86%; 7.50% of the top 1,000 |
| ProGEO.ai | Mar 2026 | Fortune 500 | not defined; automated scan, then manual review | 7.4% (37 companies) |
| Casey Burridge, from HTTP Archive | Jun 2026 | Chrome UX Report origins, by rank | 2xx and a text/plain header | 5.61% of the top 10,000 (421 of 7,504 crawled) |
| HTTP Archive Web Almanac 2025 | Jul 2025 | about 12 million desktop and 15 million mobile home pages | 2xx and a text/plain header | 2.13% desktop, 2.1% mobile |
| Common Crawl | Jul 2026 | a random sample of hosts Common Crawl fetches without problems | 200 with a text/plain or text/markdown body; redirects not followed | 11.72% |
| Ahrefs | May 2026 | 137,210 domains using Ahrefs Web Analytics | 200, Markdown rather than HTML, no error text | 28%, which Ahrefs calls "an upper bound" |
| SE Ranking | published Nov 2025 | "nearly 300,000 domains", not described | not described | 10.13% |
| llmtxt.info | Sep 2026 | a fixed panel of 219 notable hosts, mostly developer-facing | 200, plain text, starting with a Markdown H1; unreachable hosts excluded | 57.4% (124 of 216) |
| SerpPrism | Sep 2026 | 78 hand-picked large sites | 200, not HTML, final path still /llms.txt | 28.2% (22 of 78) |
| ALLMO.ai | Jan 2026 | the 50 domains Ahrefs lists as most cited in Google's AI Mode | not described | 1 of 50 |
The spread follows selection: the more a population leans towards technical, SEO-aware or developer-facing sites, the higher its rate. The publishers say so themselves: Ahrefs' customers "skew more technical and SEO-aware than the web at large", and llmtxt.info's panel "skews toward developer-facing companies".
Why the counts disagree: five choices
Every figure above comes from five decisions, and each of them moves the result.
Which sites are counted
A top-10,000 list, a vendor's customers and a panel of developer tools are three different populations. The 7.4% for the Fortune 500 and the 57.4% for llmtxt.info's panel are both accurate. Only the Common Crawl sample is random, and its authors describe the limit: "the sample is random only within the set of hosts Common Crawl fetches successfully, which is not the same thing as the web." They also topped the sample up with 27,394 hosts already known to serve the file.
What counts as a file
The tests range from a status code to a parsed document. The HTTP Archive's metric, behind the Almanac and Casey Burridge's figures, requests /llms.txt from the loaded page and records it as valid when the status is OK and the Content-Type includes text/plain. It does not look for Markdown, so a plain-text error page passes. It also rejects a file served as text/markdown, a type the proposal does not forbid: the proposal names no content type at all.
The gap between a 200 and a file is large wherever a study reports both. Thunderbit saw 1,606 responses with a 200 in the Tranco top 10,000 and 586 valid files, and wrote that a crawler counting every 200 "would overestimate valid adoption by about 2.74 times." Chris Humphrey found that 313 of 1,050 responses with a 200 were "soft 404s, serving an HTML page rather than a real Markdown file." Common Crawl: 22.32% of attempted /llms.txt URLs returned a 200, and 11.72% returned a text body. Of these studies, only Rankability sends a control request to a path that cannot exist, which catches a site that answers every URL with the same Markdown.
What goes in the denominator
Blocked, timed-out and unresolvable sites either stay in the denominator or leave it. Rankability keeps them: "Read the denominator. These percentages include every sampled domain, even when a site blocked the scan or could not be reached." Its 830 adopters in the top 10,000 are 8.3% of the list, but 4,211 domains stayed unknown: blocked, unreachable or inconclusive, and 2,010 of them because the domain did not resolve. As a share of the 5,789 with a definite result, the same 830 are 14.3%, by our arithmetic. SEOmator reports both ways, 8.6% and 11.4%.
Root, subdomains and redirects
Every crawl in the table requests the root path of one host. The proposal also allows /docs/llms.txt, and documentation often lives on its own subdomain. Crystal Carter at Wix examined over 1,400 publicly listed files and found "28% of LLMs.txt files are on subdomains" and 10% in subfolders. A root-only crawl of example.com misses both. Redirects split the studies too: Rankability and SEOmator follow them, Thunderbit and SerpPrism require the final path to still be /llms.txt, and Common Crawl records the redirect without following it.
When, and which platform switched it on
The measurements span February 2025 to September 2026, and platforms changed the defaults in between. Wix "automatically generates and maintains your llms.txt file" for its sites; Common Crawl found Wix behind 41.34% of the files it analysed. Shopify's changelog of 28 May 2026 says every store includes a default /agents.md, and "the paths /llms.txt and /llms-full.txt also point to this content by default." SEOmator attributes 18.2% of the files it found to Shopify. The Web Almanac's SEO chapter found 39.6% of the files in its July 2025 data related to one WordPress plugin, All in One SEO, and concluded: "we cannot be sure this is always a conscious act". A rate measured after these switches counts platform defaults as well as decisions.
The same effect breaks trend lines. Rankability's first reading, 0.3% of the top 1,000 in June 2025, used a different list ranked "by organic traffic and search visibility" and excluded redirects. Rankability now says: "No month-over-month growth claim is made."
Counts that do not measure adoption
Several widely quoted numbers have no population, no stated method or both:
- BuiltWith showed 9,850,012 live sites under "LLMS Text" on 29 September 2026. Neither the index behind the count nor the detection rule is published.
- NerdyData's "18,001 websites using llms.txt" counts pages whose HTML or JavaScript contains the string "llms.txt". It does not request the file. Articles quoting "951 domains in July 2025" attribute that figure to NerdyData.
- Originality.ai reported 110,607
llms.txtfiles in September 2026 among "3M+" websites, without saying which. Its headline "sites" totals equal the sum of its llms.txt, llms-full.txt and ai.txt counts, so a site with two of the files would count twice. - SISTRIX says "less than 0.005% of all websites worldwide are using the file", with no method or date for the measurement.
- Google
site:counts, which Wix used to estimate about 120,000 indexed files in May 2026, count indexed URLs and are estimates, not a census. - Directories such as directory.llmstxt.cloud and llmstxthub list sites that were submitted or collected. They count listings, not a share of anything.
One early figure also has a slip in it: a February 2025 crawl of the Majestic Million reported "0.015% ... 15 sites", but 15 of a million is 0.0015%.
llms-full.txt is not part of the proposal
llms-full.txt, one file with a site's whole documentation, is common enough to be measured: 1.03% of the Tranco top 10,000 in May (Thunderbit), 1.2% in September (Rankability) and 2.5% of the top 100,000 (SEOmator). The name does not appear in the proposal: not in the current text at llmstxt.org, and not in version 1, which mentioned only two files the FastHTML project generated from its own llms.txt, llms-ctx.txt and llms-ctx-full.txt. Version 2 dropped that passage. Mintlify, which serves llms-full.txt for its customers, says the format was "subsequently included as part of the official llms.txt proposal"; the proposal's text does not support that.
Count the file carefully. SEOmator found that half of the llms-full.txt files it read were within 5% of the size of the site's llms.txt, and Shopify's default points both paths at the same content. Our llms-full-txt check does not count a file that repeats llms.txt or redirects to the home page (how we check it).
Does anything read the file?
Adoption is not use. Google says that "Google Search itself doesn't use them" and that an llms.txt "will neither harm nor help your site's visibility or rankings in Google Search". Chrome's Lighthouse has an llms.txt audit, but marks a missing file as not applicable, "as providing the file is optional at the moment." The crawler documentation of OpenAI, Anthropic and Perplexity says nothing about reading it. Ahrefs looked at requests: of the roughly 38,000 files in its population, "97% saw no requests for it whatsoever in May."
Other uses need no crawler. Anthropic's engineering blog suggests giving Claude Code the documentation a tool depends on, noting that "LLM-friendly documentation can commonly be found in flat llms.txt files on official documentation sites". What to put in one is in our llms.txt guide. Publishing one does not guarantee that an assistant reads your site or cites it.
Which llms.txt adoption figure to quote
Quote a figure with its population and its date, never alone. For large sites, the most careful recent measurement is Rankability's: in September 2026, 8.3% of the Tranco top 10,000 domains served a confirmed llms.txt, counting every domain in the list. For the web as a whole there is no good figure yet. Common Crawl's 11.72% is the only random sample, but only of hosts it can fetch, and the Almanac's 2% is from July 2025, with a test that rejects text/markdown.
Leave out any number that cannot say which sites it measured, and any count of files without a denominator. We re-read these sources every quarter and update this post; the date at the top says when it last changed.
Our directory, counted the same way
Our scanner requests /llms.txt at the root with GoodForBotsBot and reads the body. A web page, JSON, an empty body or text with neither an H1 nor a link is not an llms.txt, whatever the status. A bot challenge or a robots.txt rule against the path counts as neither having the file nor lacking it. The figure below counts the directory's listings from their latest scans and changes as sites are scanned. The directory is not a sample of the web. What the figure does show is how far the count moves when the definition changes, on the same sites.
Live from the Good for Bots directory
What /llms.txt returns on 117 listed sites
Latest scans from Sep 30, 2026 to Oct 5, 2026. Refreshed hourly.
- Publish an llms.txt
- 7564%
- No llms.txt
- 4236%
- Serve
llms-full.txt - 3026%
- Link
.mdcopies of pages - 54%
One directory, four ways to count
Answer
/llms.txtwith HTTP 20081 of 117 (69%)
What a survey reading status codes counts, catch-all pages included.
Serve an llms.txt file
75 of 117 (64%)
Markdown with an H1 title or a link: not a web page, JSON or an empty body.
Pass our llms.txt check
63 of 117 (54%)
An H1 title, link entries and enough text to be useful.
Pass our descriptive llms.txt check
37 of 117 (32%)
A summary, and descriptions on most of the links.
Counted from each listing's latest scan with the current checks (llms-txt v2, llms-txt-quality v2 and llms-full-txt v3). The directory is not a random sample of the web. It holds sites we chose to scan and sites people submitted, so these figures describe the directory, not the web.
We look only at the root, so a site that keeps its file under /docs/ counts as having none here, as it does in every study above. The criteria for each result are on our standards page.
Frequently asked questions
What percentage of websites have an llms.txt file?
It depends on which websites. Crawls of the largest sites between May and September 2026 found a real file at the root of 5.9% to 9.3% of them. Common Crawl found a text file at 11.72% of a random sample of the hosts it can fetch in July 2026. The HTTP Archive measured about 2% of the home pages in its July 2025 crawl, with a test that only checks the status and the Content-Type header.
Is llms-full.txt part of the llms.txt standard?
No. llms.txt is a proposal published at llmstxt.org, not a standard, and it does not mention llms-full.txt. Documentation platforms serve the file; Mintlify says it developed the format with Anthropic.
Does Google use llms.txt?
Not in Search. Google's guide to its generative AI features says Google Search "doesn't use them" and that an llms.txt "will neither harm nor help your site's visibility or rankings in Google Search". Chrome's Lighthouse audits the file but treats a missing one as not applicable.
Sources
- The /llms.txt file, v2 — Jeremy Howard, llmstxt.org
- LLMS.txt adoption research report — Rankability
- llms.txt: We Scanned the Web's Top 100,000 Domains — SEOmator
- The State of llms.txt: What the Top 10,000 Sites Actually Do — Chris Humphrey
- The Rise of llms.txt: How Websites Are Signaling to AI — Thunderbit
- Signaling the Shift to Generative Engine Optimization (GEO) — ProGEO.ai
- Does anyone actually have an llms.txt? I checked millions of websites — Casey Burridge
- Web Almanac 2025: Generative AI — HTTP Archive
- Web Almanac 2025: SEO — HTTP Archive
- llms_txt_validation custom metric — HTTP Archive
- A Content Analysis of llms.txt Files from the July 2026 Crawl Archive — Common Crawl
- We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read — Ahrefs
- Does LLMs.txt impact your AI visibility and citations? — SE Ranking
- llms.txt Adoption: Who Uses It? — llmtxt.info
- llms.txt Adoption: 22 of 78 Big Sites — SerpPrism
- llms.txt for AI Search Report — ALLMO.ai
- LLMS Text — BuiltWith
- Websites using llms.txt — NerdyData
- LLMs.txt Tracking Study and Live Dashboard — Originality.ai
- llms.txt: how useful is the file for LLM optimisation? — SISTRIX
- Debunking LLMs.txt Myths — Wix Studio AI Search Lab
- Crawling a Million Websites in Search of LLMs.txt — Chris Green
- Understanding your site's llms.txt file — Wix
- Customize /llms.txt, /llms-full.txt and /agents.md — Shopify
- What is llms.txt? — Mintlify
- Writing effective tools for agents — with agents — Anthropic
- Google's guide to optimizing for generative AI features on Google Search — Google
- llms.txt audit — Chrome for Developers

