The data nobody sells: why I started taking a daily snapshot of 211,243 companies
Korea's National Tax Service runs a clean public API for business registration status. Give it a 10-digit number, get back whether the company is active, suspended or closed. It's free, it's fast, it's accurate. It also
Korea's National Tax Service runs a clean public API for business registration status. Give it a 10-digit number, get back whether the company is active, suspended or closed. It's free, it's fast, it's accurate.
It also has no memory.
Ask it today and you learn the company is closed. Ask what it was last month and there's no answer, because that question isn't in the API. There's no as_of parameter, no history endpoint, no changelog. The present tense is the only tense available.
That gap turned out to be the most interesting thing I've found while building an API for AI agents, so I did the obvious thing: I started writing down what I saw, every day.
What I built
A scheduled job, running once a day, that asks for the current status of every company in a fixed list and stores the answer. Then it diffs today against yesterday and writes a separate file containing only what changed.
Three days in, the numbers look like this:
Date Observations Changes recorded
Day 1 211,243 β (first run, everything is new)
Day 2 211,243 70
Day 3 211,243 94
The daily snapshot is about 42 MB. The changes file is 1.7 MB. That ratio is the whole point: 99.2% of what I record is confirmation that nothing happened, and the remaining 0.8% is the product.
Here's one change, as stored:
json
{"business_number":"...","field":"status","previous":"active","current":"closed",
"previous_seen_at":"2026-09-29T00:05:14Z","detected_at":"2026-09-30T00:05:14Z"}
A company that was operating when I looked on the 29th had closed by the 30th. I can't prove when it actually closed β only when I observed each state. So that's what I store: not "this company closed on the 29th", but "it was active when I looked, and closed when I looked again." Those are different claims, and for anything that might end up in a compliance report, the difference matters.
Why this is worth doing
Three reasons, in increasing order of how much I care.
One: it can't be bought. The upstream has no history. Neither does anyone else's wrapper around it. If a competitor decides tomorrow that company status history is valuable, their history starts tomorrow. Mine starts today. That gap only widens, and no amount of funding closes it.
Two: it changes the shape of the business. A status lookup is a one-shot transaction β the agent asks, pays two cents, leaves, and may never return. A watchlist is a standing relationship: register 100 suppliers, and every day I check them and tell you only what moved. Small, recurring, and from my side it runs on infrastructure that's already doing the work.
Three: change is the signal, not state. "This company is closed" is a fact. "This company suspended operations three months ago and resumed last week" is an inference about risk, and it's the kind of thing a procurement system actually wants to know. You cannot derive the second from any number of present-tense lookups. You can only accumulate it.
The things that nearly broke it
The job that ran nothing and exited 0. I wrote about this one in the last post: an entry-point guard that compared hand-built file URLs, worked on Windows, silently failed on Linux. The container started, did nothing, exited clean, and the scheduler showed green. If I'd trusted the schedule instead of running it manually first, I'd have "collected" a week of nothing.
Day one proves nothing. On the first run, every record is new, so the diff path never executes. The output looked perfect and told me nothing about whether the comparison logic worked. I only got real validation on day two, when 70 changes appeared and I could check that the previous values, current values and both timestamps were all populated. If you build something that compares today to yesterday, you don't have a working system until the day after you ship it.
Scope discipline is a security feature. The obvious way to decide which companies to track is to watch what users look up. I deliberately didn't. The corpus is built only from public registries β the procurement vendor list and the debarment list β and I enforced that with a runtime guard and a test, not just a comment, because a comment doesn't stop a future version of me who is in a hurry. Query contents never enter the corpus, which means the dataset contains nothing I learned from a customer.
*The cost
*
Roughly 2,100 upstream calls a day against a limit of 1,000,000. Storage is a few hundred won a month. The job runs in about two and a half minutes. It's the cheapest thing in the entire project and probably the most valuable, because every other part of it could be rebuilt by someone else in a week, and this one can't.
The general version
If you're wrapping a public data source for agents, it's worth asking what the source doesn't answer. Usually there's a tense problem: it tells you what is, not what was, not what changed. That absence is the only part of the data you can own.
Everything else you expose is a thinner version of something the upstream already provides. The history you accumulate is yours, and it starts the day you decide to write things down.
Which is a long way of saying: if you're going to do this, start the clock today rather than when you need it. The data you don't record is the only kind you can never get back.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.