Your AI Review Bot Reviewed a Diff the AI Itself Wrote. That's Not a Review.
The PR opens. The bot leaves a comment 12 seconds later: LGTM, looks clean, no issues found. The diff it reviewed was generated by the same model that's now reviewing it. This is the most common failure mode in AI-ass
The PR opens.
The bot leaves a comment 12 seconds later: LGTM, looks clean, no issues found.
The diff it reviewed was generated by the same model that's now reviewing it.
This is the most common failure mode in AI-assisted development right now, and almost nobody is calling it what it is: it's not a review. It's the author reading their own essay out loud to a friend who already agrees.
What an AI review bot actually sees
When you hook a model into GitHub and point it at a pull request, it reads the changed lines, the surrounding context, and your commit message. It then produces a plausibly written summary of what the code does, plus a few stylistic nits.
If the diff came from the same model (or a related one), the bot shares the same blind spots. It misses the same off-by-one. It agrees with the same questionable abstraction. It calls the same clever workaround "clean code" because it wrote the workaround.
The review loop closes. The LGTM lands. The PR merges. Three days later, production breaks in a way that was visible in the diff all along β to a human who wasn't invested in the solution being correct.
The only review signal that matters is independent
Useful code review doesn't come from another model's opinion about the model's output. It comes from something the code didn't generate itself:
- A type checker that rejects a signature the model got subtly wrong.
- A contract test that compares the actual response shape against a spec the model didn't write.
- An integration test that exercises the real path, not the mocked path the model assumed.
- A linter configured for your codebase's specific rules, not a generic style pass.
These signals don't care what the model intended. They don't read the commit message. They just run and report.
An LGTM from a model is worth roughly the confidence the model already had in its own output. Which is, by construction, higher than the evidence supports.
Where the model's review is still useful
This doesn't mean AI code review is useless. It's useful for the things a tired human reviewer skips:
- Naming conventions in new files.
- Obvious resource leaks (unclosed handles, missing await).
- Missing error branches the human might gloss over.
- Documentation drift between the code and the comment.
But it should be positioned as a second pair of eyes on style and edge cases, not as a correctness check. When the bot says LGTM, the human reviewer should treat that as "the stylistic pass is done," not as "this is safe to merge."
A cheap team policy that actually works
After watching this play out across a few repos, we landed on a rule that costs almost nothing:
Never let the bot's LGTM be the only automated signal on a PR. If the diff was AI-generated, require one independent check that the model didn't generate: a green type check, a contract test, or an integration test against a live endpoint.
For API work specifically, the independent signal is usually the contract. A spec file that existed before the agent touched the code is the one artifact the model can't rationalize. When the agent's generated handler returns a response that doesn't match the spec, the contract check fails β regardless of how confident the model sounds.
That's the reason we ended up building this into Powerduck: the spec stays the source of truth, and the agent has to re-run the request against the live endpoint and compare the actual response against the contract, not against what it thinks it should have returned. The bot can't talk its way out of a failed assertion.
The uncomfortable part
Teams that skip this step usually tell themselves they'll "just review the PR more carefully." But the whole point of using an AI coding agent is that the PR volume goes up. You can't review twice as many PRs at the same depth as before.
Which means you need the independent signal to do the work the human used to do β not another model's reassurance.
The next time an AI review bot stamps LGTM on a diff the AI wrote, ask one question: what didn't the model generate that just ran and passed? If the answer is nothing, the review didn't happen.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.