The four-line diff that charges your customers twice
Here's a change any of us could approve tomorrow. Checkout fails when the bank is slow, so you ask your agent to retry. It wraps the charge call in a retry: three attempts, on timeout. Tests pass. The diff is four lines
Here's a change any of us could approve tomorrow.
Checkout fails when the bank is slow, so you ask your agent to retry. It wraps the charge call in a retry: three attempts, on timeout. Tests pass. The diff is four lines. You approve it.
Now one question. If the bank already took the money and only the response timed out, what does the retry do?
It charges the customer again. Nothing in those four lines makes the second attempt the same charge as the first.
The code was easy to read, and you read it. You just didn't understand it. No tool in your pipeline checks for that.
Reviewing is one thing, understanding is another
Linters and tests check the code. AI reviewers leave comments on your PR. All of them ask whether the code is OK.
Nobody asks the person whose name is on the commit whether they know what it does.
That used to come for free. If you wrote it, you mostly understood it. Now an agent writes a few hundred lines, you skim, it looks reasonable, you merge. Your name is on it, and your pager goes off for it at 3am.
aye aye
"Aye aye, captain" means two things. I heard you, and I understood.
Coding agents nailed the first one. They do exactly what you ask. We're building for the second.
aye aye sees every change your agent makes, as it happens. By the time a PR exists, nobody remembers which of fifty edits actually mattered.
Most changes are low-stakes, and it lets them through without a word. A check that fires on everything becomes a rubber stamp, and a rubber stamp is how we got here.
When a change really matters, it asks you one question, out loud, in a browser tab. Something like: "A charge times out, but the bank already took the money. Walk me through what happens next."
An independent judge checks your answer against what the code actually does. The agent's summary of its own work doesn't count. The judge is also a different model from the one that wrote the code, so nothing grades its own homework.
Nobody fails. You pass, or you get walked through what you missed, and the results stay with your team.
In the retry example, passing means you can explain the change, double charge included. Then you fix it before it merges.
Where we are
The capture side is built and tested. The check ships on 29 November.
The waitlist is at ayeayecaptain.vercel.app. Move your cursor around while you're there. The logo watches you.
I'd love your take
Has an AI-written change ever passed review and still bitten you later? What question would have caught it?
And what would make a check like this annoying enough that you'd turn it off?
I read every comment.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.