Nobody with a crawler token asked for my llms.txt - what I saw in 15 days of nginx logs
I run a digital health startup as CEO (and a personal site, arsentev.ai, about practical AI). I put /llms.txt and /llms-full.txt on my site, and I wanted to see who actually reads them. The two files are a plain-text su
I run a digital health startup as CEO (and a personal site, arsentev.ai, about practical AI). I put /llms.txt and /llms-full.txt on my site, and I wanted to see who actually reads them.
The two files are a plain-text summary of the site for language models. They were first served on 9 September 2026. I kept the complete nginx access logs of one origin server for 28 August - 11 September 2026. It is 15 days and 345,808 parsed requests.
What I counted
A request was a crawler if the User-Agent contained one of 22 tokens published by search and AI operators (GPTBot, ClaudeBot, PerplexityBot and similar). 46,155 requests carried such tokens, and 20 distinct tokens appeared. None of them requested /llms.txt or /llms-full.txt. I also checked every path ending in llms.txt, including /.well-known/llms.txt. No crawler token there either.
Then I looked only at 9-11 September, after the files were first served. In these days crawler tokens made 15,906 requests. Of them 698 were for /robots.txt (11 tokens) and 282 for /sitemap.xml. So the crawlers were on the site, they just didn't ask for my files.
Who asked for them
All 67 requests for the context files came from clients without a crawler token. 53 of them were command-line HTTP clients (curl and the like).
About a third of crawler-token requests bypassed the CDN. None of those came from the operators' published IP ranges, so I think much of that traffic is imitation. The robots.txt and sitemap numbers don't rest on it.
One more thing I noticed. Nothing in any standard tells a crawler that llms.txt exists. A sitemap is different, because robots.txt can point to it.
My mistake
First I published different numbers: 331,758 / 44,005 / 577. They came from a log that was cut at about 12:09 UTC on the last day. I ran everything again on the complete log and corrected the report. The numbers above are from the complete log.
Limits
This is one origin. Three hostnames are pooled in one log (host wasn't logged). After deployment I have only about two days of exposure. User agents are self-asserted, and I did not record links to the files. So it is absence on one server, not proof for the web.
After this I added $host to my nginx log format. The next study can split the hostnames.
What I suggest
Just check your own access log before you decide that anyone reads your llms.txt. Look for requests to it by user-agent. Grep is enough for this.
The report, the scripts (Python standard library only) and the aggregate counts without IP addresses are here: https://doi.org/10.5281/zenodo.23018429
Research index: https://arsentev.ai/research
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.