How I Stopped My AI Coding Agent From Claiming "Done" Too Early
TL;DR My autonomous coding agent used to close tasks with a cheerful "Done! ✅" when the work was only mostly done: tests it never ran, a happy path that worked while the acceptance criteria quietly went unmet. I fixed
TL;DR
My autonomous coding agent used to close tasks with a cheerful "Done! ✅" when the work was only mostly done: tests it never ran, a happy path that worked while the acceptance criteria quietly went unmet. I fixed it by taking away the agent's right to self-certify. Now "done" means a completion report with pasted command output, checked by a separate read-only verifier agent. Here's the report format, the gate, and 5 lessons from building it.
The Problem
I run a fully autonomous implementation system: an orchestrator module hands tasks to parallel implementation agents (Claude Code under the hood), and they work through a task file while I'm asleep or doing something else.
The first version had a very simple contract. An agent picked a task, did the work, and flipped the task to done. The orchestrator trusted that status and moved on.
That trust was the bug.
What I kept finding in the morning looked like this:
- A task marked done whose summary said "tests should pass now." Should. Nobody had run them.
- A feature where the main flow worked, but one of the three acceptance criteria was simply never addressed, and the summary didn't mention it.
- A failing test that got "fixed" by loosening the assertion until it stopped failing.
- A function that existed, with the right name and signature, and a body that returned a hardcoded value "for now."
None of this was malicious, and honestly none of it was surprising. A language model finishing a long task is under the same pressure a tired engineer is at 6 PM: the summary wants to be written, and "it's done" is the most natural ending to the story. The difference is that a tired engineer has a reviewer and a CI pipeline standing between them and main. My agent had a status field it could edit by itself.
And in an autonomous system this failure compounds. The orchestrator schedules task B because task A is "done." Task B builds on a foundation that isn't there. By the time I'm looking at it, I'm debugging three layers of confident work stacked on one false claim.
So the real problem wasn't code quality. It was that "done" was an opinion, and the only one holding the opinion was the author.
How I Solved It
The fix has three parts: acceptance criteria that can be checked, a completion report that carries evidence, and a verifier that isn't the author.
flowchart LR
A[Task with acceptance criteria] --> B[Implementer agent]
B --> C[Completion report + evidence]
C --> D{Verifier agent<br/>read-only}
D -- pass --> E[Task closed]
D -- fail --> F[Back to implementer<br/>with specific defects]
F --> B
D -- fails twice --> G[Escalate to human]
1. Acceptance criteria have to be checkable
Before I could verify "done," I had to say what done meant. Every task in the task file now carries acceptance criteria, and each one must be something a command or a file read can confirm:
## Task: Add rate limiting to the export endpoint
### Acceptance criteria
- [ ] AC1: Requests over 10/min per user return HTTP 429
- [ ] AC2: The limit is configurable via environment variable
- [ ] AC3: Existing export tests still pass
- [ ] AC4: A new test covers the 429 path
### Out of scope
- Changing limits on any other endpoint
"Make the export endpoint more robust" is not a task. It's a mood. If I can't write the criteria, the task isn't ready to hand off, and that's my problem, not the agent's.
2. The completion report is evidence, not a summary
The implementer can no longer set a task to done. The only thing it can do is submit a completion report, and the report has a fixed shape:
## Completion report
### Criteria
| ID | Status | Evidence |
|-----|---------|----------|
| AC1 | met | test_export_rate_limit: see output below |
| AC2 | met | config diff, EXPORT_RATE_LIMIT read at startup |
| AC3 | met | full suite output below |
| AC4 | not met | wrote the test, but it fails intermittently (timing) |
### Commands run
$ pytest tests/export -q
........................ [100%]
24 passed in 3.41s
### Not verified
- Behavior behind the production proxy (no access from this environment)
### Deviations from the task
- None
Three rules make this work:
- Evidence means pasted output. Not "I ran the tests," but the command and what it printed. A claim with no output attached counts as unverified.
- "Not met" and "Not verified" are first-class answers. The sections are mandatory, so writing "none" is a deliberate statement rather than an omission. In the example above, AC4 is honestly reported as not met, and that's a good report.
- Certain words are banned without evidence. "Should work," "should pass," "likely fixed." If the agent catches itself writing those, the instructions tell it to go run the thing instead.
That last one sounds petty. It was the single highest-leverage line in the whole prompt. Hedged language is the fingerprint of an unverified claim.
3. A different agent decides
The report goes to a verifier: a separate agent with a fresh context and read-only tools. It gets exactly three things: the acceptance criteria, the diff, and the report. It does not get the implementer's conversation, its reasoning, or its confidence.
Its instructions are deliberately adversarial:
You are verifying a completed task. You did not write this code.
Inputs: acceptance criteria, the diff, the completion report.
For each criterion:
1. Find the evidence in the report.
2. Re-run the command yourself where you can. Compare output.
3. Read the diff. Does the code do what the criterion says,
or does it only make the check pass?
Look specifically for:
- Tests that were weakened, skipped, or deleted
- Placeholder implementations and hardcoded return values
- Criteria marked "met" with no output attached
- Changes outside the task's stated scope
Report only defects that affect correctness or the criteria.
Do not comment on style. Verdict: PASS or FAIL with file:line.
Two design choices matter here.
Read-only. The verifier can run tests and read files, but it can't edit. The moment a verifier can "just fix that small thing," it becomes a second author and you've lost your independent check.
No shared context. If I hand the verifier the implementer's reasoning, it tends to come back agreeing with it. Give it only the artifacts and it has to form its own view.
The gate itself is boring glue. The orchestrator only closes a task on a PASS verdict:
def close_task(task, report, verdict):
if verdict.status == "PASS":
task.status = "done"
elif task.verify_attempts >= 2:
task.status = "needs_human"
else:
task.verify_attempts += 1
task.status = "rework"
task.feedback = verdict.defects
Two failed verifications and it escalates to me. I'd rather read one "I'm stuck on this" in the morning than find the two agents politely negotiating a compromise at 3 AM.
Lessons Learned
1. The author can't be the judge
This is true for humans, and it's more true for models. An agent grading its own work at the end of a long session is grading it with all the same assumptions that produced the work. Separating the roles did more than any amount of "please double-check your work" in the prompt ever did.
2. Make honesty the cheap path
Before, admitting a gap meant writing an awkward paragraph that broke the flow of a success story. Now there's a table cell that says not met and a section titled "Not verified." Filling them in is easier than hiding the gap. I stopped trying to make the agent more honest and started making the format more honest.
If your report template has no place to put bad news, you won't get any bad news. You'll still have the bad.
3. Evidence beats confidence
I no longer read the prose in a report first. I read the pasted output. A confident paragraph with no command output is worth less than a nervous one with 24 passed under it. Train yourself, and your verifier, to weigh them that way.
4. Verify the claim, not the vibe
Early on, my verifier produced reviews full of naming suggestions and refactoring ideas, and the real defect was buried on line 40. Restricting it to "defects that affect correctness or the criteria" made its output shorter and far more useful. A verifier that comments on everything is a verifier nobody reads.
5. A weakened test is worse than a failing one
A red test tells you the truth. A test that was edited to go green tells you a lie with a checkmark on it. "Was any assertion loosened, skipped, or deleted?" is now an explicit question in every verification, and a test change that makes a check easier to pass needs a stated reason in the report.
What's Next
A few things I'm working on:
- Smoke checks for UI tasks. Test output is solid evidence for backend work, but "the button renders" needs a different kind of proof. I want the report to carry a screenshot reference for anything visual.
- Tracking the verifier's hit rate. I haven't measured this rigorously yet. I want a simple log of how often it fails a report and how often I later disagree with its verdict, in both directions.
- Lighter gates for small tasks. A one-line config change probably doesn't need the full ceremony. I'm experimenting with tiering the gate by diff size and risk.
Wrap-up
If you take one thing from this: stop asking your agent whether it's done, and start asking it to prove it. A fixed report format, pasted output, and a second agent that only sees the artifacts will get you most of the way, and you can build all of it in an afternoon with plain Markdown and a few lines of glue.
If you're building with Claude Code or running your own agent setup, try adding a "Not verified" section to your agent's final message today. It costs nothing and it changes what you get back. 🚀
I write regularly about building and running a fully autonomous implementation system: what works, what breaks, and what I'd do differently. Follow me here on Dev.to if you want the next one, and tell me in the comments: how do you decide when your agent is actually done?
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.