Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Is That Really GPTBot? Verifying AI Crawlers in Your Access Logs

Everyone writes about what to put in robots.txt for AI crawlers. Almost nobody checks whether those crawlers show up, or whether the traffic calling itself GPTBot actually comes from OpenAI. I wrote a longer version of t

Is That Really GPTBot? Verifying AI Crawlers in Your Access Logs

Everyone writes about what to put in robots.txt for AI crawlers. Almost nobody checks whether those crawlers show up, or whether the traffic calling itself GPTBot actually comes from OpenAI. I wrote a longer version of this on DevToolLab; this is the short path.

The short answer: grep your access log for the bot tokens, then compare each request's IP address with the ranges the vendor publishes. The user agent alone proves nothing, since any script can send it.

For the examples I wrote a 15-line nginx-style log for the demo. The IPs come from the vendors' real published ranges, plus a few made-up addresses acting as impostors.

Step 1: count what claims to be a bot

Each vendor puts its name inside a longer user-agent string. OpenAI's own docs show OAI-SearchBot wrapped in a full Chrome string, so an exact match on the field finds nothing. Search for the token instead:

grep -oE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|Claude-SearchBot|PerplexityBot|Perplexity-User' access.log \
  | sort | uniq -c | sort -rn

On my sample log that gives 6 GPTBot, 3 OAI-SearchBot, 3 ClaudeBot, 2 ChatGPT-User and 1 PerplexityBot. Those are claims, not facts yet. If you want a quick look without a terminal, pasting lines into the Nginx Log Analyzer shows top user agents, IPs and endpoints, but it does not verify who owns an IP.

Who publishes what

OpenAI, Anthropic and Perplexity each split crawling into separate agents and publish IP lists as JSON. As of October 9, 2026:

  • OpenAI: GPTBot (training), OAI-SearchBot (ChatGPT search), ChatGPT-User (a user asked about a page). Lists at openai.com/gptbot.json, searchbot.json and chatgpt-user.json.
  • Anthropic: ClaudeBot, Claude-User, Claude-SearchBot, all covered by one list at claude.com/crawling/bots.json.
  • Perplexity: PerplexityBot and Perplexity-User, with perplexitybot.json and perplexity-user.json.

The lists age differently. The ChatGPT-User file had 234 prefixes dated October 7, 2026, while Perplexity's two files date from February and October 2025. Refetch them on a schedule rather than copying ranges into a firewall once. OpenAI also says its bots may add a robots.txt marker to the user agent when fetching that file, which helps if your logs omit paths.

Step 2: check the IP

All five files use the same schema, so one script covers every vendor with just the standard library:

import ipaddress, json, re, sys

LISTS = {"GPTBot": "gptbot.json", "OAI-SearchBot": "searchbot.json",
         "ChatGPT-User": "chatgpt-user.json", "ClaudeBot": "claude.json",
         "PerplexityBot": "perplexitybot.json"}

def nets(path):
    return [ipaddress.ip_network(p.get("ipv4Prefix") or p["ipv6Prefix"])
            for p in json.load(open(path))["prefixes"]]

ranges = {bot: nets(f) for bot, f in LISTS.items()}
line_re = re.compile(r'^(\S+) .*"([^"]*)"$')

for line in open(sys.argv[1]):
    m = line_re.match(line)
    if not m:
        continue
    ip, ua = m.groups()
    for bot, prefixes in ranges.items():
        if bot in ua:
            ok = any(ipaddress.ip_address(ip) in n for n in prefixes)
            print(bot, ip, "verified" if ok else "SPOOFED")
            break

A log line claiming GPTBot/1.4 is checked against openai.com/gptbot.json: a match means verified GPTBot, no match means spoofed

My longer script on DevToolLab also prints a per-bot summary table. On the sample log it found 15 requests: all of the ChatGPT-User, OAI-SearchBot and PerplexityBot lines verified, 4 of 6 GPTBot lines and 2 of 3 ClaudeBot lines verified, and 3 requests came from 203.0.113.45 and 198.51.100.7, documentation-range addresses no vendor owns. Use ip in network for the test. Anthropic's list has many single-address /32 entries, and arithmetic on addresses breaks on those.

What logs will never show

Google-Extended does not appear in any log. Google's documentation says it has no separate user agent: crawling happens under existing Google agents, and the token is only a robots.txt switch for whether content may be used to train Gemini. You can set it, but you cannot measure it.

Agents that act on a user's request behave differently too. OpenAI says ChatGPT-User is not an automatic crawler and robots.txt may not apply to it, and Perplexity says the same of Perplexity-User. A hit from one after a Disallow is documented behavior.

Four mistakes to avoid

  • Matching the entire user-agent string instead of the token.
  • Counting or allowlisting a bot before its IP checks out.
  • Opting out by blocking IPs. Anthropic warns that this can stop its crawler from reading robots.txt, so use robots.txt to opt out and the IP list to verify.
  • Trusting an old copy of a vendor's list.

Where this stops

Verification tells you who sent a request, not what happens to the content afterward. Your robots.txt remains the only control the vendors document for training use, and you can check what it allows for a given bot with the Robots.txt Tester. Logs hold visitor IPs, so run the script locally rather than pasting raw logs into a hosted service.

Takeaways

Count with grep, verify every claimed bot against its vendor's list, and treat unverified ones as scrapers. If an expected crawler never shows up, check robots.txt and firewall rules before assuming it stays away. The full walkthrough, with the sample log and complete output, is on DevToolLab.

References

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.