0 of 550 top homepages send noai. 21 send Bing's noarchive.
I fetched the homepages of the Tranco top 1,000 and read every meta robots tag and every X-Robots-Tag header they sent. 550 homepages came back as real pages. None of the 550 sends noai or noimageai. None sends nosnippet
I fetched the homepages of the Tranco top 1,000 and read every meta robots tag and every X-Robots-Tag header they sent. 550 homepages came back as real pages. None of the 550 sends noai or noimageai. None sends nosnippet either, and the 47 that send max-snippet all set it to -1, which means "no limit". The page-level AI rule that big sites actually use is Bing's: 21 of the 550 send noarchive or nocache, the two values Microsoft says keep a page out of Copilot answers or cut it down to a link and a snippet.
So the AI policy of the top 1,000 lives almost entirely in robots.txt. That's fine, as long as you know what each lever reaches. Most of the confusion I see comes from mixing them up: a noai tag nobody major reads, a noindex behind a robots.txt block that stops the crawler from ever seeing it, or a header and a meta tag that say different things.
Below: what each lever reaches, the numbers, three real sites with three different setups, and how to write one policy that says the same thing everywhere.
Limits first
-
noaiandnoimageaiare not a standard. They aren't in Google's list of supported robots rules, and the crawler pages from OpenAI, Anthropic and Perplexity describe robots.txt tokens as the opt-out, not page tags. Some tools honour them. img2dataset, which builds image-text training sets, skips images whose own HTTP response carriesX-Robots-Tag: noaiornoimageai. It reads the header on the image, not a meta tag on the page around it. - Homepage only. One URL per site. A site can put directives on articles, images or PDFs and leave the homepage clean, and I wouldn't see it. "0 of 550" is about homepages, not about whole sites.
- One network location, one user agent. One GET per domain on 2026-10-06, 15:59 to 16:01 UTC, from one connection in Japan, with a desktop Chrome user agent, redirects followed. A server can send different headers to GPTBot than to Chrome. I didn't test that. I also didn't run JavaScript, so a meta tag injected client-side is invisible here (and to most crawlers).
- List and set: Tranco list 56WKN, top 1,000, the same 742 domains that answered HTTPS in my llms.txt count and my robots.txt count.
- 192 of the 742 didn't give me a homepage. 114 answered 401 or 403 to a browser user agent, 8 answered 200 with a bot wall or an error page (microsoft.com's said "Your request has been blocked"), and 70 answered 404, 400, 5xx, 429, something that wasn't HTML, or nothing. Those 192 are unknown, not "no directives". n = 550 below.
What each lever reaches
| Lever | Who reads it | What it does |
|---|---|---|
robots.txt User-agent: GPTBot etc. |
Crawlers that name that token | Whether the bot may fetch the URL at all. Details per token in the robots.txt post |
nosnippet, max-snippet:N (meta or header) |
Google: nosnippet "will also prevent the content from being used as a direct input for AI Overviews and AI Mode". max-snippet limits how much can be used |
|
noindex, none
|
Search engines | Out of the index, and so out of everything built on it |
noarchive |
Bing | Microsoft: "will not be included in Bing Chat answers, not be linked to in the answers", and not used for training. Google no longer lists it |
nocache |
Bing | Microsoft: only "URL/Snippet/Title" in answers; only those used for training |
noai, noimageai
|
Some dataset tools | Non-standard. See the limits above |
Sources: Google's robots meta tag reference and AI features page, Microsoft's Bing webmaster blog, 2023-09-22, img2dataset. All read 2026-10-07.
Two rules hold across all of these.
A crawler that robots.txt keeps out never reads your page-level rules. Google says it directly: if a page is disallowed in robots.txt, "any information about indexing or serving rules will not be found and will therefore be ignored." A noindex on a page Googlebot can't fetch does nothing.
Meta tag and header are the same instruction in two places. Google treats them as one set, and "the more restrictive rule applies". The header is the only option for PDFs and images. The meta tag is easier to change in a CMS.
The numbers
n = 550 homepages, 2026-10-06.
| Homepages | |
|---|---|
| Any meta robots tag | 192 |
Any X-Robots-Tag header |
15 |
noai or noimageai (meta or header) |
0 |
nosnippet |
0 |
max-snippet |
47, all -1 (no limit) |
noindex for every crawler |
7 |
noarchive or nocache
|
21 |
Directive aimed at a named AI crawler (GPTBot, ClaudeBot, ...) |
0 |
The 7 noindex homepages aren't AI policy. They're login apps (outlook.com and live.com land on Outlook's mail login), two domain registrars, a cloud login, an ad-tech company and a parked domain.
The 21 Bing signals split in two:
-
10 aim only at Bing:
<meta name="bingbot" content="noarchive">orX-Robots-Tag: bingbot: noarchive. instagram.com, threads.com, theguardian.com, bbc.com, bbc.co.uk, elpais.com, cnet.com and nikkei.com sendnoarchive; tumblr.com and theverge.com sendnocache. -
11 send
noarchiveto every crawler, among them linkedin.com, npr.org, lemonde.fr, dailymail.co.uk, usatoday.com and zoho.com.
max-snippet:-1 on 47 homepages is mostly SEO plugin boilerplate (it comes with max-image-preview:large and max-video-preview:-1). It says "use as much as you like". Nobody in the top 1,000 uses max-snippet to limit AI Overviews on the homepage.
Meta versus header
3 sites send rules in both places for the same crawler, and 2 of them send different sets. usatoday.com sends noarchive, nocache in the header and only old noodp, noydir in the meta tag. repubblica.it sends noarchive in the header and not in the meta tag. Neither is a contradiction, since the stricter value wins. But anyone reading only the HTML would get the policy wrong, and that includes the person who maintains the template.
10 sites have more than one meta robots tag on the homepage. I found no page that says index in one place and noindex in another.
Against robots.txt
I compared every homepage with the same site's robots.txt from the earlier run (545 files, RFC 9309 matching).
-
noaiagainst robots.txt that lets training crawlers in: 0, because there's nonoaito compare. -
A directive aimed at an AI crawler that robots.txt already keeps out: 0. The only named targets were
bingbot(10) andgooglebot-news(1). -
Page-wide
noindexthat some AI crawlers can't see: 1. sharethrough.com putsnoindexon its homepage and blocks all nine AI tokens in robots.txt. Both say no, so nothing breaks. The tag just never reaches those bots.
Of the 21 Bing-signal sites, 15 served a robots.txt file. 8 of the 15 fully block GPTBot and 10 fully block Google-Extended. 2 fully block none of the six AI tokens I compared: rubiconproject.com and zoho.com.
Three sites, three setups
bbc.com: every layer says the same thing. robots.txt fully blocks GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot. The homepage adds X-Robots-Tag: bingbot: noarchive. Bingbot can't be blocked in robots.txt without leaving Bing search, so for Copilot the page-level switch is the only lever. That's the pattern I'd copy.
rubiconproject.com: page-level only. The meta tag says noarchive to every crawler. robots.txt allows all six AI tokens. If the intent is "no generative use", it has been said only to Microsoft. If the intent was only "no cached copy", then it's a leftover that now also keeps them out of Copilot answers. I can't tell which from the outside.
theverge.com: mixed, and maybe on purpose. <meta name="bingbot" content="nocache"> (Copilot may show a link and a snippet), robots.txt blocks ClaudeBot, PerplexityBot, Google-Extended and CCBot fully, and doesn't block GPTBot at all. That reads like a policy decided per company. It could also be a deal. The file doesn't say, and that's the problem with a policy spread over three places: the next person to edit it can't either.
Write one policy
Decide per purpose first, then pick the lever the target actually reads.
-
Training. Use robots.txt tokens: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Bytespider. For Bing,
noarchive. If you also wantnoai, send it as a header on images, where the tools that honour it look. -
AI answers and search. OAI-SearchBot, Claude-SearchBot and PerplexityBot in robots.txt. For Google's AI Overviews, robots.txt can't help without leaving Google Search;
nosnippetor amax-snippetvalue is the lever. For Copilot,nocache(link and snippet only) ornoarchive(out). -
Gone completely.
noindex, and then let crawlers fetch the page so they can read it.
"No training, answers with a link only" for a news-style site, written out:
# robots.txt
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Bytespider
Disallow: /
User-agent: *
Disallow: /search
<!-- every HTML page -->
<meta name="bingbot" content="nocache">
# images and PDFs, where a meta tag can't go
location ~* \.(jpg|jpeg|png|webp|pdf)$ {
add_header X-Robots-Tag "noai, noimageai" always;
}
Then check for the four contradictions I looked for above:
- a page-level rule for a bot that robots.txt keeps out (it's never read),
- a meta tag and a header for the same crawler with different values,
- two meta robots tags that disagree,
- a page-wide
noarchiveornoindexleft over from an old SEO setup that now also decides your AI policy.
Write the policy down in a comment at the top of robots.txt, even if it's only one line. robots.txt is the one place every crawler reads, and the comment is the only part a person reads.
The raw data (one JSON line per homepage with the raw X-Robots-Tag values and meta tags, the saved <head> of each page, the probe and analysis scripts, and a control file that plants each directive to check the counts) is kept with the date above.
The free checker shows, for GPTBot, ClaudeBot and PerplexityBot, whether robots.txt allows them and what your server actually returns, with a short preview of what else it found on the page. It flags noai, noimageai, nosnippet, low max-snippet, noindex and none; it doesn't read Bing's noarchive or nocache yet.
If you want one site gone through end to end, the paid audit on the same page covers crawler access and machine-readability: robots.txt per bot, what your server answers each crawler, page-level directives, structured data, sitemap and llms.txt. It doesn't query any AI engine, so it can't tell you whether you appear in their answers. It tells you whether they can read you, and whether your rules say what you meant.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.