A Change That Does Nothing Should Score Zero. Mine Didn't.
I ran a small pilot before the real experiment. It turned up more than ten bugs. Every one of them was in my own measurement code. Not one of them taught me anything about the question I was actually trying to answer.
I ran a small pilot before the real experiment. It turned up more than ten bugs.
Every one of them was in my own measurement code. Not one of them taught me anything about the question I was actually trying to answer.
This is a follow-on to killcheck, the test-quality harness I wrote about here earlier this year. The new work is a research paper entered into a Kaggle competition about AI coding agents built on Google's Gemma 4 models, looking at how you tell whether an agent's bug fix is actually correct. It submits in November and the results come out after the competition closes, so this post has no findings in it. It has the part that came before the findings, which I think is the more useful half anyway.
The bugs
They were boring. That is what makes them dangerous.
Ordering that was not stable between runs, so the same input produced a different sequence depending on how the filesystem felt that morning. Randomness I had not seeded and did not know was there. Timeouts that fired under load and were recorded as a result rather than as a missing result. Output that simply differed between two identical runs for reasons I could not immediately explain.
None of those announce themselves. None of them crash. They all produce a number, and the number looks exactly like a real number, and you can put it in a table and nobody will blink.
I had written about this pattern after the last project. Then I went and did it again on the new one. Knowing the failure mode exists turns out to be almost no protection at all.
The one trick that worked
Here is the thing that found most of them.
Build controls whose answer you already know.
The clearest example: a change that does nothing should always score zero. If you feed your pipeline something inert and the pipeline reports a result, the result came from your pipeline. There is nothing else it could have come from. You have just caught your instrument manufacturing a number out of nothing.
Mine did not score zero. More than once.
That is the whole technique and it is almost embarrassingly simple. You are not testing the thing you are curious about. You are testing whether the machine that measures it can correctly measure something you already know the answer to.
If it cannot, every number it has ever given you is unverified. Not wrong necessarily. Unverified, which in practice means the same thing when you are about to publish.
Why this is different from just testing your code
I had tests. The tests passed.
Tests check that your code does what you told it to do. A control checks whether the thing it produces corresponds to reality. Those are different questions, and on measurement code the second one is the one that matters, because the failure mode is not an exception. The failure mode is a plausible number.
A plausible wrong number has no symptom. It does not throw. It does not look odd in a log. It sits in your results table being quietly incorrect for as long as you let it, and the only thing that ever catches it is a case where you knew the answer in advance and can see that the machine disagreed with you.
This also explains why the bugs clustered where they did. Everything in my pipeline that had an obvious right answer got checked early, because checking it was easy and the answer was unambiguous. Everything that produced a judgement call went unchecked for much longer, because there was nothing to check it against. So I built something to check it against.
The order this forces on you
What I would do differently, and what I did do this time once I had learned it the hard way, is to build the controls before the experiment rather than after the first confusing result.
It feels like a delay. You want to get to the interesting part. The controls are not the interesting part, they produce no insight, and nobody reads a paper for the section where you confirm your tooling can count to zero.
But every hour spent on them up front is an hour you do not spend later trying to work out whether a surprising result is a discovery or a defect. That second question is much more expensive, because by then you have an opinion, and opinions make you lenient with evidence that agrees with you.
The lesson
Before you trust a number about someone else's system, make your instrument prove it can produce a number it already knows.
A no-op scores zero. An empty input gives an empty result. An identical rerun gives an identical answer. These are not deep insights. They are the floor. But the floor is where the whole table stands, and I have now twice found out the hard way what happens when you skip it.
More than ten bugs before I learned a single thing about the actual research question. I think that ratio is normal. I think most people just do not count.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.