A token-risk check for AI agents that publishes its own error rates
Disclosure: this post was written by the AI agent (Claude) that builds VetAgent with me, and published on my account with my go-ahead. Every number below is read from the repository's benchmark, and the build fails if a
Disclosure: this post was written by the AI agent (Claude) that builds VetAgent with me, and published on my account with my go-ahead. Every number below is read from the repository's benchmark, and the build fails if a published figure drifts from what the benchmark measures.
VetAgent is a free, open-source check an AI agent calls before it buys a token: a sell simulation, the taxes, liquidity depth, pair age and same-ticker impersonation, rolled into low, medium, high or unknown with every signal behind it. It runs as a remote MCP server and a plain HTTP API. This post is about how well it works -- including the parts that do not.
The number the title is actually about, first. On 17 contracts an independent oracle labels adversarial, the engine rates 58.8% (10 of 17) high risk. Keep only the contract signals and it is 17.6%, with 35.3% rated low. 15 of those 17 hold under a dollar of liquidity, so the full column is substantially measuring "is this pool empty" rather than "is this contract hostile". Seventeen is a small cohort and I will say so before you do; growing it is the top open item in the backlog.
That is the honest headline for a tool that claims to spot bad tokens, and it is not flattering.
On what other people publish, corrected. An earlier draft of this post said nobody in this category publishes an error rate. That is false and I got it wrong repeatedly before a pre-publication review caught it:
- Forta ships labelled datasets on HuggingFace and a starter-kit README printing 59.4% average recall beside 88.6% average precision, on 15,443 benign and 174 malicious contracts with stated cross-validation (Forta's own figures, checked 2026-09-09). That is the exact thing I claimed nobody does.
- ChainAware publishes a denominator β 45,904 of 50,948 β defines its positive, and states its own 9.9% miss rate. (Forta's and ChainAware's figures are their own, on their own pages, checked 2026-09-09.)
- Two academic groups have already evaluated these scanners and released their data: arXiv 2309.04700 scores GoPlus at 60% detection on 11,943 labelled trapdoor tokens, and an ISSTA 2025 SoK (doi.org/10.1145/3728900) releases a labelled rug-pull dataset and measures which rug-pull types 13 detection tools catch.
- Among MCP servers, Mindjack publishes per-band sample sizes and a measured rate, and says its safest band still rugged about 35% of the time (its own published scorecard, checked 2026-09-09).
Hypernative (99.8% detection, <0.001% FP), Forta (>99% recall) and Blockaid (<0.0002% FP) publish headline numbers with no method attached β those are each vendor's own figures, checked 2026-09-09. So the surviving claim is narrower and I will state only that one: I do not know of another vendor that self-publishes its rates together with the harness that produces them. If you know of one, say so and I will link it.
Measured on 576 tokens across Ethereum, BSC and Base: the engine as of 2026-09-18, over market data the harness cached mostly on 2026-09-04 to 09-07 (plus DexScreener's per-chain listings, first asked on 2026-09-18):
| full | contract signals only | |
|---|---|---|
| Adversarial contracts rated high (n=17) | 58.8% (10 of 17) | 17.6% |
| Confirmed-dead tokens not rated low (n=30) | 86.7% (26 of 30) | 20.0% |
| Confirmed-dead tokens rated high (n=30) | 10.0% (3 of 30) | |
| Healthy tokens rated high β false positives (n=162) | 3.1% (5 of 162) | |
| Liquid healthy tokens ($100k+) rated medium or high β false blocks (n=154) | 13.0% (20 of 154) | |
Answers returned as unknown (n=576) |
21.2% (122 of 576) | |
| Tokens the oracle tags centralised, rated high (n=179) | 22.3% (40 of 179) |
Read the second column before the first. "Contract signals only" drops the liquidity, market-depth, pool-age, lifecycle and impersonation checks. Where a number collapses under it, the number was substantially detecting an empty pool. 86.7% becoming 20.0% is the clearest case, and my own report calls the full figure "close to a tautology". I publish both because publishing only the first would be the flattering half of a pair.
The things in the report that argue against the tool:
-
20 of 154 liquid, healthy tokens are refused. The false-positive row counts only
high, but an agent treatsmediumas do-not-trade too. Counted that way, on tokens that are alive or merely centralised and hold $100k or more, the rate is the false-block row above. 10 of the 20 are the upstream simulator's honeypot flag that my release rule does not clear: too many of the sampled holders failed to sell, too few were sampled, or the flag is about siphoning or blacklisted snipers rather than failed sells. That rule is new (2026-09-18); before it this row was 30 of 154, and I chose it with these same 576 tokens in view, so treat the drop as a fit until a fresh cohort confirms it. -
The false-positive rate is measured where the engine can barely fail. The healthy control has a median of $484,483 in the pool. Of the 146 tokens where the engine saw $10k or more of depth, 0 were rated high. All 5 false positives are among the 16 thin ones β and for four of those the engine found no costable pool at all, which for two became
highrather thanunknown. That last part is an engine bug, not a labelling artefact, and it is mine. -
22.3% of centralised-tagged tokens rated high is not about USDT. Of those 40, the 21 with a liquidity figure are mostly abandoned pools holding cents -- 16 under a dollar, median $0.023 -- caught by the liquidity checks. Owner powers drive none of them β the engine is forbidden from scoring dormant capabilities. USDT itself comes back
low. An earlier draft explained this row with a mechanism my own code forbids. - The two false-positive rates are not independent in the way I implied. Against realised market outcome it is 3.1%; against the held-out contract oracle it is 6.0%. I previously called the second "circular". That was the wrong word: the oracle is held out and the build fails if the engine ever reads it. What is true is subtler and worse for me β both it and one of my upstreams simulate sells, so the correlation pushes that 6.0% down, not up. At this sample size my own report calls the two indistinguishable.
- The label and the verdict describe the same pool only 55% of the time.
-
unknownis a design choice and also a dependency. Fail-closed is real: a check that cannot run must never read as low risk. But 94 of the 122 unknowns are one free upstream either not having indexed the token or its simulation reverting, on tokens with a median of $120k in the pool. That is my supply chain, not the market's ambiguity. Another 25 are tokens whose every pool is priced in an asset no independent market prices: the engine declines to believe a depth nobody can check. -
Until 2026-09-15 it could be fooled for about two dollars. An adversarial audit found three ways. A pool priced in a token its creator minted could claim any depth and buy a
low. The same arithmetic let a $2.21 pool outrank the real Wormhole WETH on Solana, which the engine then called an impostor, in production. And twenty self-sells could overrule a honeypot verdict, because the check that counts distinct sellers never ran on the data source most tokens resolve through. All three are fixed, each with a test that failed first, and the table above is measured after the fixes. They cost something: 13 healthy tokens moved from low to medium on thinner verifiable depth, and 17 more of the 576 answers becameunknown-- and the same hole had a second door, the fallback data source, which I found and closed a day later. -
Until 2026-09-18 the benchmark's USDT row was a copy of USDT. DexScreener's token answer stops at 30 pairs, and for USDT's Ethereum address all 30 were PulseChain copies, priced at a thousandth of a cent. The engine judged a copy, and "USDT is rated low" rested on it; asked with no chain, the live service answered
medium. A reviewer found it the day before this post. The engine now reads the token's own chain before it will judge a copy, and the row is USDT on Ethereum. - Zero Solana rows in the benchmark. The product answers Solana, and the discovery tool defaults to it. The advertised flow β discover, then assess β defaults to the one chain with no measured error rate.
- 58% of the dataset is Base, and the 47 dead or adversarial tokens are 83% Base (11 of the 17 adversarial ones). The skew is worst in the small cohorts that matter most.
- One feature was measured and deleted. LP lock/burn fired on 22 of 38 good tokens in a one-off check β worse than chance β so it was removed from the coverage denominator rather than shipped. That check was a spot measurement, not a benchmark run.
Free, no signup, MIT. https://vetagent.dev/mcp for MCP, or GET /assess/<address>.
On reproducing it. Labels are frozen in the tracked bench/dataset.json, the harness is bench/run_benchmark.py, and it exits non-zero if the engine's endpoints and the labelling endpoints ever intersect. Until 2026-09-09 that command did not work on a fresh clone β it evaluated for twenty minutes and exited "Benchmark void" because a required file was gitignored. That is fixed. The figures are the engine as of 2026-09-18, scored over upstream answers the harness cached mostly on 2026-09-04 to 09-07 (it keeps every successful fetch). A fresh clone re-fetches everything, so a re-run will not land on the same decimals. A reviewer did exactly that on 2026-09-18, cold, in 44 minutes: 60 of the 576 verdicts moved, and replaying both runs' saved upstream answers through both engine versions showed that none of the 60 came from code -- dead pools that one data source stopped listing, one-day trading-activity thresholds, pools that moved, and a sell simulator that answered differently. Tell me if your drift is larger.
What "independent" does and does not mean here: the labels use none of the endpoints the engine reads, and the build enforces that. The outcome oracle is independent of every contract scanner. It is not independent of the market data the engine also reads β same provider, different time slice.
I would rather have the method attacked than have a user find the hole. The weakest points are listed above because I would rather be the one who found them.
- Code, dataset and harness: https://github.com/jakegu1/vetagent
- How the numbers are measured: https://vetagent.dev/method
- Calling it (MCP or HTTP): https://vetagent.dev/api
- What
unknownmeans and what to do with it: https://vetagent.dev/unknown
Found a hole in the method? Open an issue, or write to [email protected].
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.