I let an AI build and ship a product. Its tests passed 6/6, then failed 5 of 8 real cases.
Its test suite passed six out of six. Then I pointed it at thirteen real crash reports from my own server and it got three of eight. I asked Claude to build and ship a commercial product with as little help from me as p
Its test suite passed six out of six. Then I pointed it at thirteen real crash reports from my own server and it got three of eight.
I asked Claude to build and ship a commercial product with as little help from me as possible — write the code, write the tests, set up the storefront, publish it. I'd supply the thing it couldn't synthesise: a live 234-mod Minecraft server, its logs, and its crash reports.
That division turned out to be the whole story. Here's every place the real data contradicted a confident, well-tested answer.
Synthetic fixtures measure your imagination
The first version of the crash reader recognised six failure modes and had six passing tests, one per mode. Each fixture was a small log containing exactly the string its pattern looked for.
Pointed at an actual crash-reports/ folder, it diagnosed three of eight.
The same author wrote both the question and the answer, so of course they agreed. The patterns weren't wrong, they were narrow — tuned to a tidy example of each error instead of the shape errors take when a 234-mod pack falls over. Real crashes arrived wrapped in the framework's own exceptions, or as stale-jar NoClassDefFoundErrors, or as malformed resource IDs no invented sample contained.
Rewriting the rules from thirteen real reports took it to 13/13.
I don't think this is a property of AI-written tests specifically. It's a property of tests written by whoever wrote the code, and an AI will produce a hundred of them before you've finished reading the first.
A plausible theory that moved the number by 0.00 seconds
The tool took 59 seconds on a real 3.4 MB log. The obvious suspect surfaced fast — one line of 107,445 characters, a complete HTML page some mod had fetched and logged in full.
Capping line length for matching took the runtime from 58.84s to 58.84s.
Zero change. The theory was specific, well-argued, and wrong. Stage timing found the truth immediately:
read_log 0.04s
environment 0.01s
blame_mod 0.00s
diagnose >115s <- there it is
Then per-rule timing: the first rule printed in 0.04s, the second never printed at all, because a single re.search() on one line hadn't returned.
r"(?P<mod>\S+).*is client ?-?only"
Unbounded \S+ followed by .*. On a long non-matching line the engine tries every split of the first against every split of the second. Bounding it:
r"(?P<mod>\S{1,80}) is client[ \-]?only"
58.84s to 1.97s.
There's now a test that fails loudly rather than hanging:
def test_no_pattern_has_unbounded_quantifier_after_a_capture(self):
bad = re.compile(r"\(\?P<\w+>\\S\+\)\.\*")
for rule in RULES:
for pat in rule.patterns:
self.assertIsNone(bad.search(pat), f"{rule.id}: {pat}")
Three confident wrong answers in a row
The feature that makes the tool worth anything: the framework annotates every stack frame with the jar it came from, so you can walk the trace, skip the ones belonging to the platform, and name the first third-party jar left.
The entire feature lives in the skip list. Version one blamed netty-common on about half of all crashes, because Netty sits near the top of any network-related trace. Fixed that — it blamed fmlloader. Fixed that — modlauncher.
Each was authoritative-looking and wrong, and each was caught only by running against reports whose cause I already knew.
Shared prefixes inflate every fuzzy match
A second tool validates resource IDs. Its first run on a real pack reported 169 typos:
wastelandmod:alicepack
did you mean: wastelandmod:icepick
That's not a typo, that's fuzzy matching with no idea what it's doing. The cause: comparing whole IDs, so a shared namespace contributed thirteen identical characters to every score before the distinguishing part was reached.
| compared | full ID | path only |
|---|---|---|
cooked_caned_fish vs cooked_canned_fish
|
0.984 | 0.971 |
alicepack vs icepick
|
0.905 | 0.750 |
Comparing paths alone: 169 candidates became 63. Three are real and still live in a pack thousands of people play.
An empty field is not proof of absence
This one is my favourite, because it happened to the monitoring, not the product.
Claude wrote a check to watch the storefront. It immediately reported that a published, paid product had no file attached — that anyone buying it would get a receipt and nothing to download. Alarming, specific, and it was about to send me to re-upload a file that was already there.
The product had two files. The API field being read is only populated when there's exactly one plain attachment; with two, or with one embedded in rich content, it returns {}. The check inferred "missing" from "empty".
There's now a test named after exactly that false positive, and the check reads an endpoint that actually enumerates the files.
Four 404s are not proof that an API can't
Related failure, same root. Claude told me file upload through the storefront's API was impossible and that I'd have to upload every build by hand. It had guessed four endpoint names, got 404 on all of them, and concluded.
Sending a file to the product update endpoint returns an error message that documents the entire upload flow — presign, upload parts, complete, attach. Full automation had been available the whole time.
The lesson I'd actually generalise: these APIs describe themselves when you send them something wrong. Probing only the endpoints you expect tells you about your expectations.
It named a real person's project as the thing that broke a server
The one that mattered most. The README, the test suite and a published article all used a real third-party mod as the example crash culprit — because it genuinely was, in one of my logs.
That was about to ship to paying customers as the illustration of a broken mod. It's not mine to do that to someone. It was caught by a check written to scan its own output for exactly this class of leak, and replaced everywhere with an invented name, including in the already-published article.
If you let an AI write marketing copy from your real data, the data is going to come with it.
What I'd actually tell you
Claude was fast and genuinely good at the writing. What it could not do was know whether any of it was true, and it was equally confident either way. Every correction above came from contact with a real system — not from more reasoning, more tests, or more careful prompting.
So the useful split wasn't "AI does the easy parts". It was:
- It writes. Quickly, consistently, with better test hygiene than I'd manage at 1am.
- Reality checks. Real logs, real crashes, real API responses, real users.
The second half is the one that found every bug. If you're letting an AI build something, your actual job is to be the part of the loop that touches the world — and to stay suspicious of green test suites, including the ones it writes to prove itself right.
The tools are a crash-report reader and a modpack ID validator for modded Minecraft servers. Free editions, single files, no dependencies:
The full write-up of everything it got wrong, with the fixes: jaakoby.github.io/how-this-was-built.html
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.