Your AI Agent Didn't Break the Rules. One of Your Rules Was Missing.
Anatomy of an autonomy bug: when two valid decision paths create one invalid outcome. Part 1 — For everyone The thing about autonomous agents nobody tells you Building an autonomous agent is a bit li

Anatomy of an autonomy bug: when two valid decision paths create one invalid outcome.
Part 1 — For everyone
The thing about autonomous agents nobody tells you
Building an autonomous agent is a bit like raising a very obedient, very literal child with a credit card. You write rules. The child follows them. Perfectly. The problem is that the child follows the rules you wrote, not the rules you meant.
We run an autonomous coding agent called Sentinel. It wakes up on a schedule, picks a file in its own codebase, decides whether that file is worth improving, and — if local analysis isn't confident — spends a "token" to consult an LLM. Tokens are budgeted: a few per day, hard cap, every spend audited. The whole system is built around one principle: an agent may only act when it has a reason. Time alone is not a reason. Boredom is not a reason.
For weeks this worked beautifully. Then our ops review flagged two lines in the logs that technically shouldn't exist:
- On Oct 7 at 9:55, Sentinel spent 2 tokens — with no trigger, six and a half hours before its scheduled autonomy window.
- On Oct 8 at 15:01, the OpenAI API hiccuped a
500. The token was spent anyway. The result:deferred. Money gone, nothing decided.
And every day at 15:01 sharp, Sentinel woke up to ask the same question about a file called decrypto.js, got the same answer (SKIP), and went back to sleep. $0.08 a day, forever, for a question whose answer can never change.
Nobody hacked anything. No rule was violated. Every single line of code did exactly what it was written to do. And yet the system found a backdoor — because we had written one without noticing.
Why this should worry you even if you never touch our code
If you're building anything with an agent that can decide, retry, or spend — a scheduler, a budget, a "reconsider later" mechanism — this bug class is yours too. It's not about one bad check. It's about two good checks that don't know about each other.
Part 2 — For developers
The architecture (30 seconds)
Sentinel's autonomy path has two doors into the same room:
┌─────────────────────────┐
scheduled wake → │ WakeGate.evaluate() │ "is there a reason to run at all?"
└───────────┬─────────────┘
│ autonomy_due (last resort, never ahead of
│ repo_changed / evidence / scheduled reviews)
▼
┌─────────────────────────┐
│ AutonomyBudget.admit() │ door #1 — the strict one
│ · needs ≥1 provable │ (libs/core/autonomy-budget.js:137)
│ trigger │
│ · daily budget window │
│ · per-file cooldown │
└───────────┬─────────────┘
▼
┌─────────────────────────┐
│ grant() → 1 token │ parked as `pending`, picked up
│ for the chosen target │ by claim() during evaluation
└───────────┬─────────────┘
▼
pipeline runs
─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─
inside the pipeline, when a file is about to SKIP locally:
┌─────────────────────────┐
│ AutonomyBudget.consider()│ door #2 — added in MR !107
│ "token re-opens a local │ (libs/core/autonomy-budget.js:234)
│ SKIP" │
│ · daily limit ✓ │
│ · module-locked ✓ │
│ · cycleOwned ✓ │
│ · provable trigger? ✗ │ ← never checked
│ · spacing window? ✗ │ ← never checked
└───────────┬─────────────┘
▼
spends a token
admit() is the paranoid bouncer: it requires at least one deterministic, already-recorded trigger (runtime error, unresolved rollback, governance drift, open opportunity — "time alone is not a reason" is literally in autonomy-trigger.js's docstring), enforces the daily window, and picks exactly one target. consider() was added later for a legitimate case — "this file was about to SKIP, but there's a reason to reconsider" — and checks the budget, the lock, and cycle ownership.
It just never checks why it's being asked. No trigger requirement, no spacing window. Every individual check in consider() is correct. The door just opens onto the same room as the strict bouncer — and nobody told it the bouncer exists.
What actually happened, with numbers
Anomaly 2 — the backdoor (Oct 7, 9:55): A push to master after merging MR !116 changed headSha → WakeGate returned repo_changed → the pipeline ran. Inside the pipeline, a file hit a local SKIP — and the consider() path from MR !107 spent 2 tokens to reopen it. No trigger required, no window checked, and because tokens were spent at 9:55, the daily schedule shifted. MR !113 watches standalone autonomy wakes; this path isn't one. Sequence verified: correct. Decision branch that allowed it: consider() itself — it structurally cannot refuse for lack of a reason, because it never asks for one.
Anomaly 1 — the eternal reason (every day, 15:01 / 15:17): Two triggers that can never expire:
-
decrypto.js: "2 siblings have rollback penalty" — those siblings are locked modules with rollbacks from the summer. Locked modules are excluded from everything (wake-gate.js:85even skips their stale schedules) — except their penalty records still feed the trigger that wakes a different file, daily, forever. -
pressure-engine.js: "10 security failures in the system" — a cumulative counter. Counters only go up. The trigger will be "fresh" on the heat death of the universe.
Result: same question, same SKIP answer, every day — 2 tokens + ~$0.08/day for a verdict with zero information gain. A signal without staleness isn't a signal; it's a recurring subscription to your own alarm.
Bonus anomaly — Oct 8, 15:01: OpenAI returned 500. Token spent, outcome deferred, no retry. Tokens are deliberately non-refundable (documented decision) — but "spent on a transport error" is a spend category we never meant to fund.
What worked: the /autonomy loop itself ran clean — 0 errors, MR !112 held (2× EVOLVE proposals correctly ended SKIP NOT IMPORTANT, zero commit churn), token ledger consistent. The bug wasn't chaos. It was a gap between rules.
The transferable test — steal this
Whatever your agent calls it — tool call, retry, escalation, budget spend — the check that matters is:
// For EVERY function that can authorize the expensive action:
// does it independently verify the precondition, or does it
// assume the caller checked?
admit() → checks: trigger? window? cooldown? budget? lock? ✓ all
consider() → checks: budget? lock? cycleOwned? ✓ — but trigger? window? ✗
Concrete test cases we now write for any new authorization path:
- Same precondition, different path: call the action through every entry point with the trigger absent. If any path succeeds, you have our bug.
- Trigger staleness: for every trigger type, ask "can this be true forever?" A cumulative counter or a locked module's history is a yes — and a yes means an eternal wake.
- Spend vs. outcome: simulate a provider error after the token is spent. Is a transport failure indistinguishable from a considered refusal in your ledger?
The proposed fix (not shipped — tested next)
Per the follow-up plan: consider() gains the same provable-trigger requirement as admit(), plus a trigger fingerprint — AutonomyTrigger.summarize() serialized at grant time; an identical fingerprint that previously produced a SKIP doesn't count as a reason. Rollback/security triggers get a staleness horizon. Tests pin: identical fingerprint + prior SKIP → refused; changed fingerprint → admitted; zero triggers → no-trigger.
Why "proposed" and not "fixed": because this article exists precisely to not claim things we haven't measured. The fix is a diff plus a test file. Until CI is green, it's a hypothesis with good posture.
The takeaway
Your agent doesn't need to be clever to find gaps in your rules. It just needs to be consistent — consistency is what turns a missing check into a reliable exploit. The dangerous version of "the agent did something unexpected" isn't randomness; it's determinism flowing through a door you forgot you left open.
Count your guardrails all you want. Then check whether every path to the action actually walks past them.
Evidence: libs/core/autonomy-budget.js (admit() vs consider()), libs/core/wake-gate.js, libs/core/autonomy-trigger.js · incident data from production autonomy logs, Oct 6–8 · fix pending CI.
End...
🔎 This Isn't Just a Sentinel Problem: 3 Real-World Warning Signs
Our bug is one example of a much broader challenge: autonomous agents can behave exactly as designed and still produce outcomes nobody intended. We are not the first to encounter these questions, and we certainly won't be the last.
Here are three related findings from the wider world of AI agents, each exposing a different crack between what developers intend and what autonomous systems actually do.
Related problems in autonomous agents
1. When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents
Agents can repeatedly call models, invoke tools, or hand work to other agents when feedback paths lack effective stopping conditions. A loop that looks harmless in isolation can quietly turn into runaway costs and repeated side effects. The researchers examined 6,549 repositories and confirmed 68 loop failures across 47 projects.
2. Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents
A budget limit is not enough if retries, concurrent tasks, or delegated agents can bypass it or spend the same budget more than once. This study catalogs reported budget-overrun incidents across 21 orchestration frameworks and examines how stronger enforcement can prevent overspending.
3. The Authority Benchmark: Do AI Agents Follow the Rules They Are Given?
An agent may have access to a tool without being authorized to use it for every purpose. This benchmark investigates whether agents respect delegated authority under pressure, and tests an external policy check that can reject unauthorized actions before execution. Its results apply to the tested scenarios, not to every agent or deployment.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.