← Back to blog
Mara Ellison

llms.txt Statistics 2026: What 584,107 Files Show

Common Crawl found 584,107 llms.txt files, 68% from templates. We validated a random 1,000: 75.2% fail the format. The data, and what a good file needs.

DataCrawlers
Summarize withChatGPTClaudePerplexityGrokMarkdown
On this page

The best llms.txt statistics come from Common Crawl's July 2026 crawl: 584,107 files, 68.27% of them generated from templates, with Wix alone at 41.34%. Having a file says little about its quality. We ran a random 1,000 of those files through Citlyze's llms.txt validator, and 75.2% failed the format, most often because sections list plain text instead of links.

This is the data study. For the format itself and whether to publish one, read what llms.txt is and whether it works. Every Common Crawl figure below links to its source; every validator figure is ours, with the method stated.

How many sites have an llms.txt file?

Nobody has a web-wide number. Common Crawl's 11.72% is the share of its own sampled URLs that returned a text body, and the sample included hosts already known to serve the file.

Common Crawl's content analysis of the July 2026 crawl explains the design. It added /llms.txt to the seed list for a random sample of hosts it had fetched without problems in recent crawls: 5,167,831 URLs. It also seeded 27,394 hosts that earlier crawls had found serving the file. That second group pushes the rate up, so 11.72% is not an adoption rate for the web.

What the crawl returnedFigure
URLs with a recorded outcome6,563,125
Returned 40469.8%
Returned 20019.61%, and fewer than half of those (45.63%) carried a text body
Successful responses served as HTML, typically a single-page app's catch-all519,484 (40.36%)
Successful responses that were robots.txt content at the llms.txt path136,578 (10.61%)
Plain-text or Markdown responses that were empty14,191
Files analyzed584,107

Common Crawl sets its figure beside the 2% or so the Web Almanac reported for 2025 and Ahrefs' 28%, which Ahrefs calls an upper bound, and says the populations differ too much to compare. The analysis covers one crawl, so it shows no trend.

For a trend, Originality.ai has tracked more than 3 million sites for a year: 4,088 served an llms.txt file in June 2025 and 36,120 in May 2026, an 8.8x rise, according to PPC Land's report on the study. Even after that growth, the files sat on about 1% of the sites it tracks.

The 40% of "successful" responses that were HTML pages matters if you check your own site. A request for /llms.txt that returns your home page with a 200 status looks like a file to a careless check. It is a soft 404.

Who writes llms.txt files?

Plugins and site builders wrote most of them. Common Crawl found that "68.27% of the corpus is templated", and that the ten largest template clusters alone cover 38.71% of files.

GeneratorFilesShare of the 584,107
Wix241,44841.34%
Unknown197,61233.8%
All in One SEO73,13612.5%
Yoast40,5096.9%
GoDaddy domain parking14,7112.5%
Rank Math12,3602.1%
GitBook2,4980.4%

Common Crawl published the Wix share and the file counts; we calculated the other shares from those counts. Its full report describes the detection: a self-declared signature such as a "Generated by" line, a structural fingerprint for builders that do not sign their files, and hashing of each file's template skeleton.

The same analysis found the file drifting from its original purpose. 44.87% of files mention the Model Context Protocol, the bulk of them because of Wix's template. 6.59% carry usage or rights policy language the specification never mentions. 2.54% are domains advertising themselves for sale. And 3,793 files contain mild attempts to steer AI models; 10 reached Common Crawl's top severity grade, and it judged 4 of those genuine.

Only 49.90% carry the complete shape the specification asks for: a title, a summary and sections of links. That measure checks whether the parts exist; the validator results below check every line, which is why they come out lower.

How many llms.txt files pass validation?

About a quarter. Of the 1,000 files we sampled, 11.9% passed Citlyze's validator with no warnings, 12.9% passed with warnings, and 75.2% failed. More than half of the failures come from one template.

Method. We drew a seeded random sample from Common Crawl's published dataset of the July 2026 crawl, which holds the 598,298 plain-text and Markdown responses behind the analysis above. We kept files fetched from the root /llms.txt path and skipped /llms-full.txt files, files at other paths and empty bodies, drawing again in their place, until we had 1,000 (1,069 draws in all). We then ran each file through the same validation code that powers our free llms.txt validator, with the same 64 KB limit, on October 10, 2026. The validator checks a file against the llmstxt.org format line by line and gives each file a verdict; it has no numeric score. The files are the July 2026 versions, and sites may have changed them since.

VerdictShare of files95% intervalWith cosmetic deviations accepted
Valid11.9%10.0% to 14.1%12.1%
Valid with warnings12.9%11.0% to 15.1%20.9%
Invalid75.2%72.4% to 77.8%67.0%

The last column is a second pass that also accepts a dash instead of a colon before a link's description. It shows how much of the failure rate is cosmetic: 8.2 percentage points.

Why files fail

The link-list check causes all but 14 of the 752 failures. The specification says each item in an H2 section is a Markdown link, - [name](url), with an optional colon and notes after it. Files fail when a section holds something else:

What the failing lines look likeFiles affectedWhat it means
Plain text bullets inside a link section60.4%A section used as prose, such as Wix's list of MCP tool parameters
A dash or pipe instead of a colon before the notes11.9%Cosmetic; most parsers would still read the link
A link title wrapped onto a second line2.2%Breaks line-based parsers
A URL containing an unencoded space0.4%The link breaks for any client that splits on spaces
Other malformed link lines3.1%Mixed shapes

A file can appear in more than one row. Horizontal rules (---) and plain lists placed before the first section are allowed by the specification and do not count against a file.

Results by generator

We tagged each sampled file by its self-declared generator line or Wix's structural fingerprint. This catches fewer templates than Common Crawl's full method, so "Other" includes undetected templates. The shares we detected (Wix 42.6%, All in One SEO 12.4%, Yoast 6.6%) sit close to Common Crawl's corpus shares, a useful check that the sample is representative.

GeneratorFiles in sampleValidValid with warningsInvalidInvalid, cosmetic deviations accepted
Wix4260.0%0.0%100.0%100.0%
All in One SEO1240.0%9.7%90.3%25.8%
Yoast6672.7%22.7%4.5%4.5%
Rank Math190.0%89.5%10.5%10.5%
Other36519.5%23.3%57.3%56.7%

The generator decides the verdict. Every Wix file in the sample failed, because Wix's template documents its MCP tools as plain bullets inside link sections. All in One SEO files use a dash before most link descriptions, so they fail the strict check and most pass once that is accepted. Yoast and Rank Math files follow the format in most cases. Those two groups are small, 66 and 19 files, so read their percentages as a direction rather than a measurement.

Beyond the verdict, the other checks show what the files contain:

CheckResult in the sample
H1 title on the first line83.8%
Blockquote summary after the title68.8%
At least one H2 section97.6%
No links at all26.1%
Over 20,000 characters, our size guideline14.4%
Headings deeper than H250.8%
Median size3,305 characters, 3 links, 5 sections

What does a good llms.txt file contain?

A title, a one- or two-sentence summary, and short sections of links to your most useful pages, each with a note on what the page covers. The llmstxt.org specification makes only the H1 title mandatory. Everything else is convention.

A compact example for a fictional scheduling product:

# HarborRoster

> HarborRoster is shift-scheduling software for restaurants and retail teams, with payroll export on the Teams plan.

Pricing and plan limits change each quarter; the pricing page is the source of truth.

## Product

- [Features](https://harborroster.example/features): Scheduling, shift swaps and manager approvals
- [Pricing](https://harborroster.example/pricing): Plans, limits and what each includes

## Docs

- [Payroll export](https://harborroster.example/docs/payroll-export): CSV export of approved hours, Teams plan only
- [Shift approvals](https://harborroster.example/docs/approvals): How managers approve or reject changes

## Optional

- [Changelog](https://harborroster.example/changelog): Release notes by month

To build or repair one:

  1. Open with an H1 title on the first line, using your site or product name.
  2. Add a blockquote summary of one or two sentences on the next line.
  3. Put any free-form notes before the first H2. The specification allows paragraphs and lists there.
  4. Write each section as links only. Every bullet is - [name](url): note, with a colon before the note.
  5. Use absolute URLs with no spaces, and keep each link on one line.
  6. Move skippable links to an ## Optional section so a reader with little context space can drop them.
  7. Stay under about 20,000 characters. A file that lists every URL on the site gives a reader no shortcut.
  8. Validate after every plugin update. A generator can change its template without notice.

If a plugin wrote your file, open it once. The free llms.txt validator checks a live URL or pasted text and shows the line behind each failure.

Does llms.txt affect AI citations?

There is no evidence that it does. Google's guide to generative AI features, last updated July 10, 2026, says you do not need AI text files to appear in Google Search, and that creating them "will neither harm nor help" because Google Search ignores them.

Server logs point the same way. An Ahrefs study of 137,000 domains found that 97% of llms.txt files received no requests at all in May 2026, and that AI retrieval bots made 1.1% of the requests that did arrive, as PPC Land reported. We found no statement from OpenAI, Anthropic, Perplexity or Google, as of October 10, 2026, that their crawlers read the file; our llms.txt explainer covers the log experiments site owners have run.

Chrome now checks the file: Lighthouse's Agentic Browsing category includes an llms.txt audit, one of the seven checks in our agentic SEO audit guide. Its documentation flags server errors at the path, and a site with no file gets a not-applicable result. That is a tooling signal, not proof that any answer engine uses the file.

So treat llms.txt as low-cost housekeeping. If you publish one, make it valid and current. If you do not, nothing in this data says you are missing citations. Crawler access matters more: the AI crawler guide covers which bots to allow, and the GEO audit guide puts llms.txt in order among the other checks.

Download the data

The aggregate tables from our sample are available as a CSV: llms-txt-validator-sample-2026-10.csv. It holds the verdict shares with intervals, the per-check results, the failure categories and the per-generator results. Common Crawl's raw files are on Hugging Face.

Run your own file through the free llms.txt validator to see which of these groups it falls into.