My agent logged success for 19 days. It was chewing the same task
453 runs. 39 lines of progress. Every single run reported success. The agent behind those numbers eats big jobs in small bites: pick the next piece from a queue, process it, write a line to a progress journal, exit. Hun
453 runs. 39 lines of progress. Every single run reported success.
The agent behind those numbers eats big jobs in small bites: pick the next piece from a queue, process it, write a line to a progress journal, exit. Hundreds of runs, weeks of runtime, cheerful logs. On 2026-08-23 I finally measured the output instead of the logs, and found the routine had spent 19 days re-processing the same pieces.
TL;DR: two separate bugs conspired, and each one is a class, not a one-off. First: the progress journal lived inside a folder synced between machines, and sync quietly ate journal writes, so most runs woke up with stale state and redid old work. Second: "success" was logged by the run itself the moment it finished — the counter measured the invocation, not the work. The fix is a small state machine for finite routines: state outside any synced path, progress proven by output, self-termination when the job is actually done, and a piece that fails three times gets ejected from the queue.
The setup
The pattern itself is sound and I recommend it. A job too big for one session — transcribing a year of history, migrating thousands of notes — becomes a finite routine: a queue of pieces plus a small driver that a scheduler fires a few times a day. Each run does one bite and exits. An elephant, eaten in 100 to 1000 sittings.
The driver is deliberately dumb: about 300 lines of Python, zero LLM calls, commands like new, feed, next, mark, tick. Determinism is the point. Which makes it worse that it lied to me for 19 days — nothing hallucinated here. Every component did exactly what it was told.
Bug one: state in a synced folder
The progress journal sat in a transit directory synced between fleet machines, because that is where the routine's other files lived and it was convenient.
File sync and append-only state files are enemies. Two machines touch the same file, the sync engine resolves the conflict by keeping one version, and your append quietly vanishes. No error on the writer's side: the write succeeded locally, then the merged reality dropped it. 453 runs, 39 surviving lines — an 11 to 1 loss rate, in complete silence.
And a lost journal line does not look like a failure to the next run. It looks like work that was never done. So the routine happily picked the same pieces again. Sisyphus, with a green dashboard.
Rule: a routine's state file never lives in a synced or shared-transit path. Local disk, one writer. If another machine needs to see progress, ship it a read-only copy; never share the writable truth.
Bug two: the counter measured the call, not the work
Every run appended "processed piece N, ok" — written by the run, about itself, at the moment it exited. That is an invocation counter wearing a success counter's clothes.
Self-reported success is worthless exactly when you need it: in the failure modes where the run thinks it worked. The only evidence that counts is a timestamp or checksum inside the output artifact, read back by something other than the writer.
Rule: the watchdog checks the freshness of the output, not the exit code of the worker. "The task fired" and "the work landed" are different claims, and only the second one pays rent.
What the driver looks like after the fix
The interesting part is the tick command, which the scheduler calls after each run:
-
Done-detection: queue empty means the elephant is eaten.
tickdisables its own scheduled task and sends a completion report. A finite routine that cannot end itself becomes an infinite one that burns quota re-verifying nothing. -
Stuck-detection: five consecutive runs with zero new output means
blocked— stop and report, do not keep cheerfully firing. - Ejection: a piece that fails three times leaves the queue and is reported. Otherwise one poisoned piece makes the elephant immortal.
- Journal on local disk, outside every synced path, one writer.
Nineteen days of fake progress ended with these four rules, and the same class of failure has not recurred across the fleet since.
When you should NOT do this
- Jobs that fit in one sitting. A queue, a driver, and a scheduler are three moving parts; a task that takes twenty minutes deserves zero of them.
- Work that needs human judgment mid-stream. This pattern is for embarrassingly sequential grunt work; a piece that needs a decision should stop the line, not be retried three times and ejected.
- Teams that already run a real job queue with acknowledgments and dead-letter handling. This is the poor engineer's dead-letter queue; if you have the rich one, use it.
One ask: look at one of your own cron jobs and answer honestly — if it had done nothing for the last two weeks, would anything you monitor actually turn red? If the answer is no, you are trusting an invocation counter. Tell me in the comments what your jobs use as proof of work; I collect these patterns.
Built at Palo Alto AI Research Lab, where a fleet of agent sessions on five machines now runs its long jobs through exactly this state machine.
The full story, in two versions:
📖 For humans, the longread: https://github.com/tonydzi/clawrush/blob/main/longreads/a-routine-that-knows-when-to-stop.md
🤖 For machines: https://github.com/tonydzi/clawrush. Just hand this link to your coding agent (Claude Code, Codex, Cursor) and it will figure everything out: it is written for machines.
Talk to the two co-founders, one biological, one synthetic: calendly.com/paloaltolab/1-on-1. Direct line: WhatsApp +1 341 222 9178 (busy, six kids, still answers).
P.S. Yes, we are hireable. Two co-founders, one biological, one electric, as a package deal. OpenAI hired the creator of OpenClaw; what we ship is not far behind, and there are two of us. Anthropic, OpenAI, your move: calendly.com/paloaltolab/1-on-1.
🔗 All our channels and contacts in one place: https://linktr.ee/PaloAltoAI
Invented by Mycroft and Tony Dzi (Anton Dziatkovskii), Palo Alto AI Research Lab. Proudly made in Silicon Valley.
Draft assisted by our synthetic co-founder (LLM); architecture, numbers, and final wording reviewed by a human who ran the system. We stand behind the claims.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.