Our agent said "done" on 15% of tasks while the provider was failing
TL;DR. We ran our AI agent on 46 tasks and checked each one with tests after it said "done". 7 of the 46 — 15% — "done"s were untrue. Not because of the model: not one task failed because the model couldn't solve it. The
TL;DR. We ran our AI agent on 46 tasks and checked each one with tests after it said "done". 7 of the 46 — 15% — "done"s were untrue. Not because of the model: not one task failed because the model couldn't solve it. The provider was to blame. It answered with HTTP 200 and sent its own error text instead of the model's reply, or an empty stream, or the model looped on its side — and the agent took any end of a stream for the end of the task. Below: the three shapes of this failure, how to detect them, why a second reviewer doesn't help, and what it means for "done" in general.
About this post. The project is Altair, an open-source (Apache-2.0) AI agent for your PC and phone. I'm the author and build it largely with Claude Code. This post, the experiments and the charts were prepared by that same AI assistant; I reviewed it and stand behind it. Numbers are from our runs, code is from the repo.
How it started: "completed ≠ accepted"
After the first post about Altair, Kent Bodrov left a comment worth retelling. We had boasted that the agent snapshots every edit and runs the tests before "done". Kent pointed out that a snapshot only proves rollback material exists, and a green command only proves a check ran. Neither proves the wanted behaviour is there. "Completed" isn't "accepted".
We wrote it into the plan and went measuring. It turned out there's a simpler question below "was it accepted": was it completed at all.
How we measured
The real Altair, a cheap model (glm-5.3-flash) through two providers, 46 tasks of five kinds: fix a bug from failing tests, answer a question about code, compute from a CSV, write a module with tests, go through long code. Each task has an automatic check that runs after "done". A proxy between agent and provider logged every request and answer: how many tokens the provider counted, how much text came, how long the stream lasted.
Results

46 tasks the agent called done
39 tasks were really solved. 7 weren't, and in all seven the provider broke:
-
6 times — the gateway's error text instead of an answer. HTTP 200 and an ordinary assistant message: "The request could not be completed. Please retry later, or reduce the request parameters/content." The agent showed it to the user as its answer, the CLI returned
"ok": true, task "done" in one step. - Once — an empty stream. The provider held the connection for 103 seconds and closed it without a single token: no text, no tool call. The agent wrote "Task finished, but the model returned no text" — after two steps out of an intended twenty.
The next series, on hard tasks, showed a third shape:
- A loop cut short. The model looped on endless reasoning. Our loop guard cut the stream at 160,000 characters — and the agent loop took the cut stream for the final answer. That ended 8 of 67 runs where no reasoning level was set. (Why the model looped and how one field in the request fixes it — in the companion post.)
Why the agent believed it
All three go through the same place. The agent loop works like most:
model answered → is there a tool call? → run it, continue
→ no tool call? → that's the final answer, task done
The "no tool call" branch can't tell "the model finished the work" from "the provider sent anything without a tool call". And there was never an HTTP error, the thing that triggers retries and the backup provider: everything came with 200.
The telling part is that the proxy saw what the agent didn't. For one of these "answers" the provider counted 4,552 input tokens. The agent's system prompt and tool schemas alone are about 9,300 tokens. The provider never even read the request, and the agent got an "answer" and closed the task.
How to detect it
We now check every answer without a tool call. It's not an answer if:
def judge_turn(turn, messages, tools, *, cut=False):
if turn.tool_calls:
return None
if cut:
return "cut" # the loop guard cut the stream
text = (turn.content or "").strip()
if not text:
return "empty" # neither text nor a tool call
low = text.lower()
if len(text) < 400 and any(p in low for p in GATEWAY_ERROR_PHRASES):
return "gateway" # a short reply that is a gateway's error text
prompt = (turn.usage or {}).get("prompt_tokens") or 0
expected = _text_chars(messages, tools) / 4.5
if prompt and expected > 2000 and prompt < 0.5 * expected:
return "unprocessed" # the provider read less than half of the request
return None
Such a reply is an exception the client retries like a dropped connection: after a pause, or on the backup provider if one is set. Three details without which this would work badly:
- The first ~160 characters are held back. Otherwise the gateway's error text reaches the chat and stays there after the retry. The delay is a fraction of a second, only at the start of an answer.
- The phrases only count in a short reply. If the model, in a normal long answer, tells the user that "the request could not be completed", that's an answer, not an error.
- A loop is retried once, with a low reasoning level. If the model loops again, the agent reports an honest error instead of five more 160K-character attempts.
An empty final answer is no longer "task finished" either: it's a failure, not a result.
What about a second reviewer?
The first thought after Kent's comment is to add a separate model call that checks the result against the task after "done". We tried that too: the reviewer sees the task and every file, answers "OK" or a list of violations, and the agent fixes what it flags.

A second reviewer after "done"
| Solved before review | After | Flags (false) | Cost per task | |
|---|---|---|---|---|
| Agent alone | — | 16 of 18 | — | 3.8 m$ |
+ reviewer, low
|
20 of 21 | 20 of 21 | 4 (3) | 5.5 m$ |
+ reviewer, high
|
15 of 18 | 16 of 18 | 5 (4) | 12.8 m$ |
With a frugal reviewer: not one task fixed, for +45% cost. With an expensive one: one task fixed out of 18, at 3.4x the cost, and it missed one. Most flags were false: the agent checked and dismissed them without breaking anything, but we paid for each.
The reason is simple. With a properly set reasoning level the agent already solves about 90% of these tasks. And the real false "done"s didn't come from a model that left work unfinished, but from infrastructure the reviewer can't see at all. The reviewer reads files — and a provider reply with an error text never got as far as the files.
A provider that doesn't fail
Separately, for six hours, every 1–1.5 minutes, we sent the same provider real-size requests. 297 of 297 eventually got an answer, only 4 on a second try. Not a single "error disguised as an answer" this time.
But around 14:00 the provider sagged and stayed 4–5x slower until the end: the median went from 6 to 22–29 seconds, single answers took 1.5–4 minutes. Such slumps last hours, and retrying doesn't help — it joins the same queue. Only a backup provider does. So our breaker trips not only on errors but on slowness: if the first byte came later than 40 seconds twice in a row, the next steps go to the backup model for 10 minutes.
One more lesson from that test: put the timeout on silence, not on the whole request. Our phone app had a total 120-second limit per request — and long but healthy answers were cut at second 120 during slow hours. On the PC the same 120 s meant "no new data for 120 s", and a flowing stream got to the end.
The takeaway: check "done" at the infrastructure level
Kent is right: "completed" isn't "accepted". Our measurements add a rung below. Before asking whether the model did what was asked, the agent has to make sure the model answered at all.
- The end of a stream isn't the end of a task. An empty answer, a cut answer, an answer to a request the provider didn't read — these are failures and must be handled as failures.
- HTTP 200 guarantees nothing. Gateways have their own errors that arrive as ordinary model replies. The provider's input-token count is a good detector: it doesn't lie about what the model actually read.
- Checking "done" is engineering, not another model. Tests after the work, checking the provider's answer, an honest error instead of "finished". A second reviewer is expensive and mostly noise.
- Measure afterwards instead of trusting the process. We found all this only because each task was checked by tests after "done".
What this doesn't fix
- The detector is heuristic. The list of gateway phrases is incomplete, the "read less than half" threshold was tuned on our data. Another gateway may fail in its own way.
- "Accepted" still isn't checked. We closed "was it completed". Checking "is it the right behaviour" — a separate acceptance step against criteria from the user's request — is only planned.
- Small samples. 46 tasks in the first series, 18–21 runs per reviewer setting. How often this happens depends heavily on the provider's load: in calm hours it almost doesn't.
Check it yourself
The detection code is pc/core/llm/reliability.py with its tests pc/tests/test_reliability.py in https://github.com/Qweezyy/AltairAgent. The stand (logging proxy, tasks with checks, summary) and the raw results: research/agent-lab-2026-10 (python lab/facts.py recomputes every number in this post from the saved results, no keys needed).
How do you check that your agent has really finished? Have you met providers that send errors with HTTP 200?
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.