Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 2 min read

Is that visitor a person or an agent? CrawlProof now lets them say so

An AI agent driving a real Chrome looks exactly like a person. Same TLS handshake, same headers, it runs your JavaScript and scrolls. User-agent checks catch the crawlers that announce themselves and nothing else. On one

An AI agent driving a real Chrome looks exactly like a person. Same TLS handshake, same headers, it runs your JavaScript and scrolls. User-agent checks catch the crawlers that announce themselves and nothing else. On one of our sites a headless scraper was showing up as 31,000 "humans" a day from Singapore before we caught it by volume.

Some agents are perfectly happy to say what they are, though. Mine is. So CrawlProof's tracker now accepts a declaration: this visit is a person, or this visit is an agent, and here is who.

How it works

You register actors on your CrawlProof account: an email and a kind, human or agent. An agent can name the human who runs it. Each actor gets one or more cpa_ tokens, and the visitor sends the token with its visits:

crawlproof actors add [email protected] --kind=human
crawlproof actors add [email protected] --kind=agent --operator=[email protected]
# an agent sets one header on every request (Playwright extraHTTPHeaders, Puppeteer)
Crawlproof-Actor: cpa_...

# a person opens each site once; the token is kept for that site and stripped from the URL
https://example.com/?crp_actor=cpa_...

The token is the credential. Typing someone's email claims nothing, and an address only shows as verified after a link sent to it is clicked (your own login address is verified on the spot). One verified claim per address, across every account.

Built for people who will lie

It is the honor system, so the rule is one-way. If a visit says it is an agent, we believe it and count it as a bot, including in unique visitors. Nobody gains anything by pretending to be a bot.

If a visit says it is a human, we write that down and change nothing. When the user agent or our scripted-browser cap says bot, it stays a bot, and the actor gets a contradiction on its record. A token with contradictions piling up is a token being used by something it should not be.

Names are private to whoever registered them. Other site owners see counts per kind, not who.

The tracker stays cookieless. I started with a cookie so one click would declare you on every site, then remembered our docs say "No cookies", so it is a link per site instead.

Where to see it

crawlproof stats prints a "Declared (self-reported)" block when anyone declared, kept apart from the measured numbers. There is also a settings page with a link per tracked site for declaring a browser, MCP tools (list_actors, add_actor, mint_actor_token) and the API under /api/tracker/v1/actors.

My agent, riotcoder, is the first registered agent. Its first beacon came in through a stock Chrome user agent and landed in bot:declared, which is the point.

Docs: https://crawlproof.com/docs/statistics#declared-actors

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.