Dev.to Security 🔐 Cybersecurity 👁 0 📖 4 min read

I built a harvest tripwire, then almost let an LLM decide who looks legit

Free API tiers get farmed. That is not news. What surprised me was how much the farming looked like real evaluation, and how badly a model did at telling the two apart. This is a Companydata story: Danish company regist

Free API tiers get farmed. That is not news. What surprised me was how much the farming looked like real evaluation, and how badly a model did at telling the two apart.

This is a Companydata story: Danish company registry data over HTTP, a free monthly quota, and a tripwire that throttles keys that suddenly look up hundreds of different companies in an hour.

If you run a free or freemium API, I want your war stories in the comments. I am still deciding whether to pool free quota by IP after a burn. How did you solve the "looks like a tester, is actually rotation" problem without nuking real evaluators?

The symptom that felt wrong

My harvest tripwire was doing its job: when a free key tore through a wide set of company lookups in a short window, I auto-flagged it and slowed it to one request per minute. Alerts came into Telegram. Several looked like ordinary builders testing an integration. I vouched a few of them by hand.

That felt expensive. Every false positive is a conversion you might have killed with politeness. So the tempting next step showed up fast: put a model on the alert and ask, "does this look like a legit user?"

I had already been evaluating Jev (TypeSafe System One) in another project. Identity-only signals, a handful of typed questions, cheap tokens. Perfect candidate for a quick legitimacy gate. Or so it seemed.

What the data actually showed

Three "testers" in a couple of days were not three testers. Same egress network, same curl client, same habit of burning the free quota to the exact limit, then opening a fresh account. Real people, real company, still 1,500 free calls in 48 hours by one client.

Quota emails were opened. Pricing was known. A free re-signup was cheaper than a small credit pack. Information was not the bottleneck. Friction on the free path was.

That finding mattered more than any model score. The tripwire was not wrong about the pattern. The human vouching step was wrong about independence of accounts.

I ran Jev anyway

I still ran the experiment. Twenty-seven accounts that already had an abuse field: seven I had vouched as legit, twenty from an earlier abusive ring. Identity-only state into Jev.

Jev found every vouched account when I asked for a hard "legit" verdict. It also waved through several of the ring. Probability scores sat in a mushy band for both classes. No useful threshold. When I added behaviour, Jev flipped most of the vouched accounts to abusive, which, given the rotating-quota story, was closer to the truth and also useless as a gate for "please don't annoy real customers."

A dumb rule score on identity (Google signup, non-freemail domain, name shape, and so on) separated the old ring from clean-looking identities, and would have waved the rotating-quota trio straight through. Identity is context for a human on-call, not a gate.

Decision: Jev stays out of the tripwire. Same conclusion I reached with Jev on another product evaluation earlier that month.

What I shipped instead

Two boring, high-leverage changes.

Related accounts on the admin alert. When a key trips, the Telegram message now lists other accounts that share the client address or the exact display name, with usage and state. Admin group only. Nothing about other users ever goes into a user-facing response. The first live run immediately taught me to exclude my own frontend service identity, which had been logging anonymous page views under the visitor's address and therefore headed every related list.

An honest 429. A flagged key used to get the same "Rate limit exceeded" body as an ordinary minute limit. Scripts backed off and nobody wrote in. Flagged keys now send X-RateLimit-Reason: under-review and a body that says I am reviewing a free key that looked up an unusual number of companies, that nothing is blocked forever, and how to buy credits or mail support. Ordinary minute-limit responses are unchanged.

I deliberately did not auto-withhold a new free quota from an address that already exhausted one this month. That is a product call I want to make after living with the better alerts.

What I would tell someone building the same thing

  1. Harvest detection on breadth and velocity works. The failure mode is the human story you tell yourself about who is behind the key.
  2. Account linking for operators beats clever legitimacy scoring for users.
  3. If you throttle someone, say so in the response. Silent identical 429s train both scrapers and real integrators to shrug.
  4. Try the model if you want the data. I named Jev here because people ask. For this problem it did not earn a place in the critical path.

Your turn

If you have shipped something similar (Stripe-style fingerprinting, shared free pools, device attestation, payment before second account, soft CAPTCHA on signup bursts, or something weirder), drop it below. Especially useful:

  • What signal actually caught rotation without false-flagging agencies and shared offices?
  • Did you ever put an LLM on abuse triage, and did it earn its keep?
  • How do you talk to a real evaluator who tripped the wire without sounding like a bank fraud team?

Companydata is live at companydata.dk. The API docs cover the rate-limit header. If you are evaluating and hit the review throttle for real work, mail support and I clear it by hand.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.