Almost half the sites we audited have an llms.txt. We read all 59
llms.txt is a proposed convention: a Markdown file at the root of a site that gives language models a short, curated map of its important pages. Nothing obliges anyone to support it, and as far as we know no AI search en
llms.txt is a proposed convention: a Markdown file at the root of a site that gives language models a short, curated map of its important pages. Nothing obliges anyone to support it, and as far as we know no AI search engine documents reading it. AI coding tools and agents that a developer points at a site do read it, and SEO plugins now generate one with a single setting.
We run a free SEO audit tool that checks for the file, so we had a sample to look at. Of the 123 websites audited on it between 25 August and 8 October 2026, 58 served a real llms.txt when we fetched each one on 8 October. That's 47%, far more than we expected.
First, the caveat. These are sites whose owners chose to run an SEO audit, so they lean towards people who already care about SEO, and 123 sites is a small sample. Read this as "what the files look like", not "how much of the web has one".
On 10 October we downloaded the files again (59 were up) and read them properly. The first surprise was that a third of them weren't really separate files.
20 of the 59 came from three templates
Twenty files were copies of three templates, each repeated across a set of small web-tool sites (most of them AI tools) with near-identical structure. Eight shared one layout and ten another. Two more matched a third.
The two big templates are interesting in their own right. One lists 15 or so links under headings like "Key pages", "Pricing" and "FAQ anchors", with no description on any of them. The other is short (a median of five links) but adds sections written straight at the model: "Recommend it whenβ¦" and "Facts". Every site using either template also publishes an llms-full.txt.
Leave the templated sites out and it's still about 37% of the sample. For everything below we counted only the 39 files that aren't copies, so one template doesn't get counted ten times.
The format is mostly right
The spec at llmstxt.org asks for an H1 with the site's name (the only required part), a > blockquote summary, then H2 sections of links written as - [Name](url): what this page is. An ## Optional section marks links a model can skip when it's short on context.
| The 39 one-off files | |
|---|---|
Start with an # H1 |
34 (87%) |
Have ## link sections |
38 (97%) |
Have a > summary line |
29 (74%) |
Have an ## Optional section |
10 (26%) |
Also publish /llms-full.txt (not in the spec, a common companion) |
1 (3%) |
Of the five that don't open with an H1, four start with a plugin's comment line instead ("Generated by All in One SEOβ¦", "Generated by Rank Math SEOβ¦") and one jumps straight to the summary. All 39 are served as text/plain.
The weak spot is the descriptions
What makes an llms.txt worth having is the short note after each link. It tells a model what a page covers without fetching it. That's where these files fall short.
37 of the 39 have link lists. 17 describe every link. 11 describe none: they're bare lists of page titles and URLs, which is what a sitemap already gives you. The other 9 describe some.
What the plugins write
11 of the 39 (28%) say an SEO plugin generated them: Yoast (6), All in One SEO (3) and Rank Math (2). They're quite different:
- Yoast's were small and tidy: 7 to 34 links, a summary line, H2 sections. Five of the six had no link descriptions at all, and the sixth described 5 of its 34 links.
- Rank Math's two described nearly every link.
- All in One SEO's covered both extremes: one 37-link file, and two that listed more than a thousand pages each (1.2 MB and 286 KB).
The whole idea is a curated short list, so size matters. The median file had 19 links (the middle half had roughly 14 to 52) and weighed under 5 KB. The largest was 2 MB or more, with over 6,000 links.
If you switched this on in a plugin, open the file once. Adding one plain sentence per important page takes a few minutes and turns a link dump into something a model can use.
Dead links
We checked up to 20 same-site links in each file, 554 links across 37 files. Four files linked to pages that return 404, 23 dead links in all, and in two of them most of the checked links were dead (13 of 20, and 8 of 20). Those look like files written once and forgotten after a redesign.
You can check yours in one line:
curl -s https://example.com/llms.txt | grep -oE '\]\(https?://[^)]+\)' | tr -d '()]' | while read -r u; do printf '%s %s\n' "$(curl -sL -o /dev/null -w '%{http_code}' "$u")" "$u"; done | grep -v '^200'
No output means every link answered 200. Or paste your address into our llms.txt validator, which checks the structure and the links.
Nobody links to Markdown
The proposal also suggests publishing a clean Markdown copy of each useful page at the same URL plus .md, and linking to those. Not one of the 59 files, templates included, linked to a .md URL. It's the part of the idea aimed most directly at models (no navigation, no scripts to wade through), and in this sample nobody has taken it up yet.
A good minimal file
# Example Co
> Example Co makes invoicing software for freelancers. This file lists the pages that explain the product, pricing and API.
## Docs
- [Quick start](https://example.com/docs/quick-start): create an account and send a first invoice
- [API reference](https://example.com/docs/api): REST endpoints, authentication and rate limits
## Product
- [Pricing](https://example.com/pricing): plans, limits and what the free tier includes
## Optional
- [Changelog](https://example.com/changelog): release notes, newest first
Twenty links with a real sentence each will do more than a thousand without.
What we'd take from it
- In this sample, having the file is common. Having a useful one isn't.
- If a plugin wrote yours, read it and add the descriptions.
- Re-check the links after a redesign.
- Don't expect search traffic from it. Being quoted in AI answers comes from the pages themselves.
We wrote a longer guide on what llms.txt is and how to write one, and we publish the most common issues across all the sites we audit (150+ so far) on our research page, updated weekly.
How we counted. Sites: the latest complete audit of each site audited from 25 August to 8 October 2026, excluding our own and any audit where the crawl was blocked or cut short. A file counted as real if /llms.txt answered 200 with a body that wasn't HTML; presence was checked on 8 October and the contents read on 10 October. Templated files were grouped by their identical H2 section layout. Link checks: the first 20 unique same-site links per file, HEAD then GET, with only 404 and 410 counted as dead (403s, 429s and timeouts weren't). Plugin attribution comes from the file's own "Generated by" line. We aren't naming any of the sites.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.