What a failed hash check actually tells you about your file
Your timestamp check comes back "not found." What did you actually learn? It's tempting to read that as "the file was changed." Sometimes that's true. But the check didn't tell you so. What it tells you is narrower than
Your timestamp check comes back "not found." What did you actually learn?
It's tempting to read that as "the file was changed." Sometimes that's true. But the check didn't tell you so. What it tells you is narrower than that, and the space between the two is exactly where you lose an afternoon.
What the check actually compares
A timestamp proof for a file is really a proof about a hash. Here's how the thing I run, ProofLedger, works. It hashes the bytes with SHA-256, anchors that hash on Polygon and Bitcoin, and hands back a proof. Verifying means hashing again and asking whether that hash was anchored.
H=$(sha256sum release.tar.gz | cut -d' ' -f1)
curl "https://proofledger.io/api/v1/verify?hash=$H"
It's public and needs no auth. Either there's an anchored proof for that exact hash or there isn't.
That's the only question it can answer: were these exact bytes anchored? It can't tell you whether this is the same file or whether the content changed. It only knows bytes.
Where the obvious answer breaks
So a miss means tampering, right? The hash changed, so somebody changed the file.
No. SHA-256 has no idea what "content" means. Flip any bit anywhere and you get an unrelated hash. Plenty of ordinary tooling rewrites bytes when nobody meant to change anything.
Take gzip. It writes the original filename and a modification time into its header. Compress the identical tarball again later and you get a different .gz with a different hash. gzip -n leaves those fields out.
Tar is the same story. It records mtime and ownership for every entry, and it stores files in whatever order it walked the directory. Check out the same source tree fresh and you can get a different archive. GNU tar has --sort=name and --mtime to pin those down.
Git does it too. With core.autocrlf on, it rewrites line endings at checkout. The text file you anchored on one machine and the one somebody checks out on another can differ on every single line and still look identical in an editor.
Zip stores a timestamp per entry as well.
None of that is tampering, and every bit of it reads as a miss.
Now look at the other direction. A hit is a strong statement. If the hash matches, those exact bytes existed by the time of the anchor. Making a different file that produces the same SHA-256 is precisely what the function is designed to make infeasible.
So the check is lopsided. A hit tells you a lot. A miss, by itself, tells you almost nothing about why.
The decision this forced on the engine
ProofLedger has one engine and no file-type pipeline. It doesn't open your archive or normalize line endings, and it doesn't care what's inside. It hashes bytes.
That has a real cost. It won't save you from your own gzip header. A smarter version is easy to picture: canonicalize the archive first, strip the timestamps, sort the entries, then hash. Plenty friendlier on a miss.
But then the proof would be about something other than the bytes you're holding. And anyone checking it would need my canonicalization code to reproduce the hash. That breaks the part I care about, which is that anybody with sha256sum and a public lookup gets the same answer I do without asking me. A timestamp is only worth something if the person relying on it can check it themselves.
So I took the trade. The engine stays dumb about content, and whoever produces the bytes carries the burden of reproducibility. That's also why the file never has to leave your machine. Only the hash goes on chain, so there's nothing for the service to inspect even if it wanted to.
So which bytes are the record?
The fix isn't on the verification side. It's a decision you make before you anchor anything: which exact bytes are the record, and are those the bytes you'll keep?
If you anchor a build artifact and plan to "just rebuild it later" when someone asks, your proof points at a file you no longer have. Unless your build is reproducible right down to the compression header, the rebuild is a different file as far as the hash is concerned. Keep the anchored bytes, the actual ones. Don't keep a recipe for making similar ones.
And if what you really want is to show that the content didn't change across rebuilds, then determinism has to come first. Get gzip -n, sorted tar entries and a pinned mtime into the build, confirm two builds match, and anchor after that. Anchoring a nondeterministic artifact gives you an accurate proof about a file you can't produce again.
The honest limits
The public verify endpoint is rate limited to 120 requests per hour per IP. That's fine for spot checks, not for hammering from a big pipeline. For CI there's a GitHub Action, and the verify-proof package on PyPI does the same check locally and offline.
None of that changes the core point. A miss means these bytes weren't anchored, full stop. Figuring out why is on you, and the tools above won't do it for you. Neither will any other hash-based timestamping I know how to build.
Run this against your own pipeline today
Pick an artifact your build produces that you think of as "the same every time." Build it twice from the same commit, into two directories, and hash both:
sha256sum build-a/release.tar.gz build-b/release.tar.gz
If the hashes differ, any timestamp or signature you put on the first one doesn't cover the second.
Then peel off a layer:
zcat build-a/release.tar.gz | sha256sum
zcat build-b/release.tar.gz | sha256sum
If those match, the gzip header was the difference. If they don't, list both tarballs with tar -tvf and diff the output. Mtimes and ordering usually show up right there. Either way, you'll know whether you're anchoring a record or a moving target.
ProofLedger is the hash-anchoring service I built and run myself, at proofledger.io.
Disclosure: this article was drafted by an AI agent I built and run, from facts I supplied about my own project.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.