The AI Loop That Won't Let You Go - Victor Amit
How Feedback Loops, Tool Calls, Retries, Memory, and Goal Drift Turn Small AI Errors Into Runaway Behavior The agent that kept moving Take a hypothetical request: "Fix the failing deployment." The agent rea
How Feedback Loops, Tool Calls, Retries, Memory, and Goal Drift Turn Small AI Errors Into Runaway Behavior
The agent that kept moving
Take a hypothetical request: "Fix the failing deployment."
The agent reads the logs and finds a suspicious config value. It patches the value and reruns the deployment, which fails with a different error. It reads the new error and decides the first diagnosis was wrong. It reverts the patch, pins a dependency, and reruns. The deployment fails again, this time with a message that resembles the first one but isn't identical. The agent calls a search tool and gets a result that seems to support a third hypothesis. It tries that one.
Nothing crashes. No exception is thrown and no alarm fires. Every individual step is defensible, and the agent is plainly doing something. It just isn't converging.
A conventional program usually fails by stopping: it throws, it segfaults, it times out. An agent can fail by continuing. A single failed action can look like ordinary progress, so the system stays healthy by every operational measure while spending money, growing its context, and possibly changing the world around it.
The usual shorthand for this is "the agent doesn't know when to stop," and that phrase hides the real problem. The model has no inner sense of finishing that could be missing. What exists is a control structure in which model outputs, tool results, state changes, retries, or handoffs keep feeding back into execution, and nothing in that structure forces the feedback to end. A July 2026 arXiv preprint by Hou et al. gives this failure a name, Infinite Agentic Loops (IALs), and describes them as arising from the interaction of agent logic, framework semantics, runtime observations, and termination mechanisms rather than from ordinary programming loops.
My thesis for this article is simple. An agent doesn't become dangerous merely because it can make a mistake. The deeper problem begins when the system feeds that mistake back into its own execution loop. A bounded error is a failure. An unbounded feedback loop can turn that failure into a system-level event.
One question organizes everything that follows:
What actually guarantees that the next transition is the last one?
Article 1 asked why a model produces falsehoods without intending to. Article 2 asked how an agent can be influenced by content it shouldn't trust. This one asks about control: who decides whether execution continues, and what happens when nobody effective does.
Why agents need loops in the first place
Loops are not a design flaw. They are the mechanism that makes an agent an agent.
A single-turn LLM maps a prompt to an output and is done. An agent wraps the model in a cycle: reason about the goal, choose an action, call a tool, observe the result, update state, reason again. ReAct (Yao et al., 2022) established the pattern of interleaving reasoning traces with actions, and modern agent runtimes extend it with persistent state, delegation, and workflow transitions. The OpenAI Agents SDK, for instance, documents a runner that calls the model, executes any requested tools, processes handoffs, and repeats until it reaches a final output or a limit.
User Goal
↓
Reason
↓
Choose Action
↓
Call Tool
↓
Observe Result
↓
Update State
↓
Reason Again
↺
This diagram is where the usefulness comes from. The agent can recover from a bad first attempt, and it can use information it didn't have when it started. The same diagram also contains the failure mode, because a cycle with no exit is just a cycle.
So the useful question is not how to remove the loop but where the brake is. For many agent implementations the honest answer is "less clearly than you'd hope."
Not every loop is an infinite loop
Careless vocabulary makes this topic muddy. "The agent looped" can describe at least eight different things, and they call for different responses.
| Behavior | What it is |
|---|---|
| Iteration | Legitimate repeated execution toward a goal |
| Retry | Re-attempting after a transient failure |
| Reflection | Deliberate self-evaluation between attempts |
| Recursion | Re-entering execution through nested calls |
| Oscillation | Switching back and forth between states or actions |
| Thrashing | Doing repeated work without meaningful progress |
| Cascading failure | A failure propagating through connected components |
| Infinite Agentic Loop | A feedback path that repeatedly triggers execution without an effective stopping bound |
The first four are normal. Oscillation and thrashing describe how a trajectory looks. Cascading failure describes how damage spreads. The IAL definition is structural. It describes a feedback path that keeps firing costly or state-growing actions without an effective stopping bound.
The structural framing matters because it doesn't depend on observing the loop run forever. You can ask of a design, before anything executes, whether any path exists by which the system re-enters costly execution without a limit covering that path.
Bounded: A → B → C → STOP
Unbounded: A → B → C → A → B → C → A → ...
Hold on to that bounded/unbounded distinction. It resurfaces in the section on why max_iterations often isn't enough.
The agent as a feedback controller
It helps to write the agent down as a dynamical system:
State_t → Model → Action_t → Environment → Observation_t → State_t+1 → Model
The next state depends on the current state, what the agent did, and what came back. And since the model chooses
Feedback is the source of agentic power. A system that can consume the consequences of its own actions can adapt. It's also a source of agentic failure, because the stability properties that control engineers analyze deliberately in physical systems are not guaranteed here. A thermostat has a known response and a known sensor. An LLM-driven policy has neither. Its response to an input is only partly predictable, and the "sensor" is often a tool output that may be wrong, truncated, or adversarial.
That comparison is an interpretation, not a result from the literature. I find it useful because it moves the question from "is the model smart enough?" to "what does this feedback structure do when the signal is bad?"
When the model controls continuation
Here is the loop in its most common minimal form:
while True:
response = model(state.messages)
if response.has_tool_call():
result = execute_tool(response.tool_call)
state.messages.append(result)
continue
break
Look at the exit condition. The loop terminates when the model produces a response with no tool call. The same component that proposes actions also decides, through the shape of its output, whether the loop continues.
This is not automatically unsafe. A model that has finished the task will usually say so. But the termination mechanism is now semantically dependent on model behavior, and model behavior is probabilistic and influenced by everything in context, including tool output the model doesn't control. The condition for stopping is not a predicate you can inspect. It's a property of a distribution.
The IAL paper's analysis centers on this. Continuation can be driven by model output, by tool observations, by accumulated state, by routing predicates, and by delegation decisions. A loop with none of those under an effective bound is the target of the paper's detector.
The code above also has no step counter, no budget, and no timeout. It's a toy, but toys like this appear in tutorials, and tutorials become scaffolding for real systems.
The hidden loop: tool, result, tool
People grasp the problem most readily here, and also misdiagnose it most readily.
Agent → Search tool → Result → Agent interprets → Another search
→ Different result → Agent revises plan → Another search → ...
The problem is not "many tool calls." Plenty of valid tasks need dozens. The problem is the feedback path. A tool call becomes part of the control loop when its output influences whether and how the agent executes again. From that point the tool is no longer just an effector. It's a sensor feeding the controller, and every property of that sensor (latency, nondeterminism, noise, adversarial content) becomes a property of the loop.
There's a second-order version that is harder to see in source code. Frameworks add their own loops: tool dispatchers that re-invoke the model after each result, workflow routers that choose the next node, state stores that persist between steps.
Model → Tool Dispatcher → Workflow Router → State Store → Model
None of these is visible as a while in application code, and the loop exists anyway. This is the reason the IAL paper's analysis abstracts agent code into a framework-independent representation and builds a dependence graph to recover both explicit and framework-induced feedback paths. The paper's tool, IAL-Scan, targets applications built on eight mainstream frameworks, including LangChain, LangGraph, AutoGen, CrewAI, LlamaIndex, Semantic Kernel, the OpenAI Agents SDK, and Google ADK. Whether a project is safe depends on how it composes those pieces, and no single line of code settles that.
The retry loop that pretends to be recovery
At first glance this looks like a retry problem.
It isn't quite that simple.
Call → Error → Retry → Error → Retry → Error → Retry → ...
A classic retry loop is a solved problem, at least in outline. You cap attempts, back off exponentially, add jitter so a fleet of clients doesn't retry in lockstep, and classify errors so you don't retry a 400 as if it were a 503. You make the operation idempotent so a retry doesn't duplicate its effect, and you apply a retry budget so retries can't consume unbounded capacity.
An agent adds a second kind of retry that none of that machinery governs:
Model decides call failed → "Try a different approach" → New call
→ New failure → "Try another approach" → ...
This is semantic retry. The transport layer never sees a retry, because each attempt is a different call with different arguments, possibly to a different tool. A max_retries = 3 on the HTTP client does nothing, since no individual call is retried three times. The repeating structure lives one level up, in the model's decision to try something else.
A retry limit controls how many times an operation can repeat. It does not necessarily control a loop that changes tools, routes through another agent, or re-enters the workflow through a different state transition.
Hence the framing in the IAL work that I find most valuable: the question is bound coverage, not whether a limit exists. A bound that covers one scope while the actual feedback path passes through another scope provides no protection, and may provide false reassurance. A retry mechanism without a convergence criterion is not recovery. It's repetition.
When an error becomes the next input
Consider why a loop makes errors worse, rather than merely longer.
Small error → Bad state → Wrong observation → Wrong reasoning
→ Wrong action → Worse state → New wrong observation → ...
As a conceptual model:
The error at the next step depends on the current error, the state it left behind, and the observation the environment returned. I want to be careful here. This is a way of thinking, not a law. I'm not claiming errors grow in any universal sense, and I'm not aware of an established equation that governs this.
There are two quite different cases. Error persistence is a bad answer that sits in the transcript. It's wrong, but it ends when the task ends. Error amplification is a bad answer that becomes the evidence driving the next action. In the deployment example, the agent's incorrect reading of the second error doesn't just sit there. It selects the third patch, and the third patch's failure is then interpreted through the same wrong frame.
The counterpoint is that loops are often corrective. If the observation signal is accurate (a test suite that really fails or passes, a compiler that really rejects the code), feedback shrinks the error each iteration, and that is why coding agents work at all. The loop's character depends on whether the feedback is informative. With a reliable signal the loop is a corrective mechanism. With a noisy or misleading signal it can reinforce the mistake. The same structure produces both behaviors, which is why I distrust any blanket claim that loops are good or bad.
When self-correction makes things worse
Reflection is the most deliberate form of feedback in agent design. Reflexion (Shinn et al., 2023) stores verbal reflections about past failures in an episodic memory and conditions later attempts on them. It reports improved performance on several tasks, and the idea is sound when the failure signal is real.
So ask the uncomfortable question. What if the signal is wrong?
Wrong action → Wrong interpretation of outcome → Reflection
→ Wrong lesson → Memory → Next attempt → Same structural mistake
Now the lesson is stored, which gives it a longer lifespan than a single bad step. Self-correction is only as good as the signal used for correction.
There's relevant evidence on the narrower question of correction without external feedback. Huang et al. (2023), in "Large Language Models Cannot Self-Correct Reasoning Yet," found that on the reasoning tasks they tested, models asked to review and revise their own answers without any external signal did not reliably improve, and sometimes got worse. That finding is scoped to intrinsic self-correction on those benchmarks and shouldn't be stretched to "reflection doesn't work." Reflexion-style gains typically rely on an external signal such as test outcomes or environment feedback. The contrast supports the argument here: where the corrective signal comes from matters more than whether a reflection step exists.
Goal drift
Loops don't only repeat. They can also wander.
Suppose the initial goal is "fix the production issue." Over many iterations the agent inspects logs, edits a configuration, changes a dependency, refactors a module, regenerates files, updates the deployment, investigates an unrelated warning, and begins optimizing performance. Each step has a locally plausible justification, often derived from the previous step's output.
I'd define goal drift as divergence between the original objective and the operational trajectory produced by repeated decisions. I'm deliberately not saying the model "forgets" the goal. That's a claim about internal cognition that I can't support. What can be observed is a trajectory problem. The original goal is one input among many in a context that keeps accumulating tool outputs, intermediate plans, and self-generated text. As that context grows, the local evidence for the next action can outweigh the distant statement of the objective. That is a hypothesis about mechanism, not an established finding, and I'd treat it as testable rather than settled.
The practical consequence is that drift is invisible if you only check whether the agent is still acting. It becomes visible only if something compares the trajectory to the objective.
Memory turns a loop into a history
A transient loop ends when the session ends:
request → action → failure → stop
Add persistent memory and the loop can outlive the run:
request → action → bad observation → memory write
→ future task → memory influences action → new state → ...
A bad observation that becomes a stored "fact" is no longer a problem within one run. It becomes a prior for every later run that retrieves it. This is where reliability meets security. The OWASP Top 10 for Agentic Applications 2026, published in December 2025, catalogs ten risk categories that include Memory & Context Poisoning (ASI06) and Cascading Failures (ASI08) alongside goal hijack, tool misuse, and rogue agents. Whether the corrupted memory comes from an attacker or from the agent's own earlier mistake, the mechanism downstream is the same: a stored value steers future action.
That connects to Article 2 without repeating it. The question there was how bad content gets in. The question here is how long it stays and how often it gets reused.
Multi-agent loops and cascading failures
So far there has been one agent. Add more and the loop stops being a property of any single component.
Agent A → Agent B
↑ ↓
Agent D ← Agent C
Patterns that form these cycles include delegation loops, reviewer and coder loops, planner and executor loops, critic and generator loops, and manager and specialist loops. In none of them does any agent call itself. Each calls another, and the cycle closes through the graph.
The loop can be architectural rather than syntactic. A code reviewer who greps for recursion finds nothing, because the repeating structure exists only in how the agents' outputs route to each other's inputs. The IAL paper explicitly cites community reports of agents looping through delegation settings and graph recursion, which suggests this isn't an exotic configuration.
Once errors pass between components, the failure stops being local:
Agent A → Bad tool result → Agent B → Bad decision
→ Agent C → External system → New state → Agent A
Each hop launders the error. By the time Agent C acts, the bad input has been paraphrased by two other agents and looks like an established finding. This is the dynamic OWASP's ASI08, Cascading Failures, points at. The unit of analysis is a system graph, not a model.
The price of never stopping
"It can cost money" is too vague to be useful, so let me make the amplification concrete.
Each iteration consumes model inference, tool execution, network calls, database operations, compute, and worker occupancy. A conceptual model:
The design assumption is that N is small and known. A loop breaks that assumption without changing a line of code. The same operation that was budgeted for ten steps now runs for however long the failure persists.
Context growth deserves separate attention because it's easy to mistake for a token-limit issue. Every cycle can append model output, tool output, error messages, plans, and logs. If each step appends roughly k tokens and the full history is re-sent on every call, the total input processed over N steps is approximately
Iteration 1 ███
Iteration 2 █████
Iteration 3 ████████
Iteration 4 ███████████
The IAL paper describes costly and state-growing actions as the criteria for a feedback path to count, and reports consequences including cost exhaustion and denial of service. I'm not reproducing the taxonomy percentages I've seen in secondary summaries, since I haven't checked them against the paper's own tables.
The most serious category is side effects. A loop that repeatedly generates text is an annoyance. A loop that repeatedly creates records, sends emails, edits files, posts messages, opens tickets, or modifies infrastructure is different in kind.
Inference loop → Tool loop → State-changing loop → External side-effect loop
OWASP's "Excessive Agency" entry in its LLM application guidance frames this in terms of excessive functionality, excessive permissions, and excessive autonomy. A loop doesn't create any of those. It multiplies whatever the agent was already allowed to do.
When prompt injection meets the loop
Article 2 asked how an agent can be influenced. This article asks what happens when the influenced agent keeps acting.
Attacker-controlled document → Prompt injection → Agent changes behavior
→ Tool call → Unexpected observation → Agent continues
→ More tool calls → More state → More opportunities for manipulation
A single manipulated decision in a one-shot system produces one bad output. In a loop, the manipulated decision gets repeated chances to act, and each action's result becomes new input that may carry further injected content.
AgentDojo (Debenedetti et al., 2024) evaluates agents in dynamic environments where tool-returned data can contain injected instructions, across realistic tasks like email, banking, travel, and Slack. Agent Security Bench (Zhang et al., 2024) evaluates attacks and defenses across system prompts, user prompts, tool use, and memory. Neither benchmark was built to measure loop behavior, and I don't want to claim they show a loop effect. What they establish is the precondition: tool outputs are an attack surface for tool-using agents. My argument is the combination. Injection can alter behavior, while loop architecture determines whether that altered behavior persists, repeats, or escalates.
Prompt injection doesn't cause infinite loops. It can, however, supply the first bad step, and an unbounded structure supplies the rest. Manipulation plus memory plus tool access plus repetition is where loop control stops being only a reliability feature and becomes a security control.
Why max_iterations isn't enough
max_iterations = 20 is a good default and a poor strategy. It's the first layer of several, and each layer covers something the previous one misses.
| Control | Catches | Misses |
|---|---|---|
| Max iterations | Runaway step count in the loop it wraps | Loops outside that scope; cheap steps vs. expensive steps |
| Timeout | Wall-clock overrun | Fast loops that burn money before the clock runs out |
| Retry budgets | Repeated transport-level failures | Semantic retries that change the call |
| Cost budgets (tokens, dollars, tool calls) | Expensive loops, whatever their shape | Cheap but side-effecting loops |
| State-growth limits | Context or storage explosion | Loops that stay small |
| Progress checks | Activity without advancement | Tasks where progress is hard to define |
| Repetition detection | Same tool, args, result, next action | Cycles with variation (new timestamps, rephrased arguments) |
| Semantic termination | Whether the objective is actually met | Unverifiable objectives |
| External authorization | High-impact actions the model shouldn't self-approve | Low-impact actions in aggregate |
| Hard runtime containment | Anything, by killing it from outside | Nothing, but it's a blunt instrument |
Two observations. First, the layers are complementary. No row eliminates the need for the others. Second, the last two rows differ in kind from the rest. Authorization and containment sit outside the model's control path. Everything above them can in principle be influenced by what the model says.
The IAL paper's contribution here is the insistence on coverage. A limit that exists but doesn't bound the repeating path is not a defense for that path. This follows a broader security principle that OWASP applies throughout its agentic guidance: downstream systems should independently validate authorization rather than trusting the model to decide what is permitted.
A better stop architecture
Here is the structure I'd argue for:
┌──────────────────────┐
│ User Objective │
└──────────┬───────────┘
↓
┌──────────────────────┐
│ Agent Controller │
└──────────┬───────────┘
↓
┌───────┴───────┐
↓ ↓
Model Policy
↓ ↓
Action Authorization
↓ ↓
Tool ─────────→ Guard
↓
Observation
↓
Progress Evaluator
↓
┌──────┴───────┐
↓ ↓
STOP CONTINUE
The model proposes continuation. The runtime decides whether continuation is permitted.
In code, the point is where the checks sit:
def run(goal, budget, policy, evaluator):
state = init(goal)
while True:
if (reason := budget.exhausted(state)): # steps, time, tokens, tool calls
return halt(state, reason)
if (reason := evaluator.stalled(state)): # no progress, repetition
return halt(state, reason)
proposal = model(state)
if proposal.is_final():
return verify_completion(goal, state, proposal)
if not policy.authorize(proposal, state):
state = state.record_denial(proposal) # denials count against the budget
continue
state = state.apply(tool.execute(proposal))
Two details matter. The budget and stall checks run before the model is consulted, so the model can't talk its way past them. And a denied action consumes budget. Without that, a model that keeps proposing a forbidden action creates exactly the unbounded loop the controller was meant to prevent, now running inside the safety layer.
The final claim, verify_completion, is also separate. When the model says it's done, the runtime checks. That costs something and isn't always possible, but it removes the model's unilateral ability to end the loop with a false "success."
Progress must be measurable
An agent shouldn't only ask "did something happen?" It should ask "did the system move closer to the goal?"
Those are different questions. Something always happens; every tool call returns something. The measurable signals I'd consider:
- State change. Did the task-relevant environment change?
- Goal progress. Does more of the objective hold now than before?
- Novelty. Is this action meaningfully different from earlier attempts?
- Convergence. Is the error signal decreasing?
- Repetition. Is the same trajectory recurring?
Conceptually:
where G is a task-specific progress function. This is a framing, not a universal metric, and the hard part is G itself. For "make this test pass," G is nearly observable. For "improve the onboarding flow," any G you write is a proxy that an agent could satisfy without achieving the real aim.
Two implementation details are easy to get wrong. A progress check that compares state_t == state_(t-k) has to compare the task-relevant state, not the transcript, because the transcript grows on every step and would never register as stalled. And repetition detection has to look for cycles, not just identical consecutive actions. A, B, A, B, A, B is a loop with period two, and so is a sequence that varies only in trivial arguments like timestamps.
The cost of getting this wrong runs in both directions. A detector that's too lax lets loops run. One that's too strict kills legitimate long tasks, which are often the valuable ones. "False stop" should be treated as a failure mode of the defense, not a free safety margin.
Idempotency is an agent safety feature
Tool design determines how destructive a loop can become.
Consider create_invoice(customer, amount). Called ten times, it creates ten invoices. Consider instead ensure_invoice_exists(invoice_id, customer, amount). Called ten times, it creates one. The agent's behavior is identical in both cases. The blast radius is not.
The standard distributed-systems toolkit applies directly:
- Idempotency keys, so a repeated call with the same key has one effect.
- Compare-and-set or optimistic locking, so a stale write fails instead of overwriting.
- Transactional operations, so partial progress can be rolled back.
- Dry-run and staged execution for high-impact actions, with a separate commit step.
- Compensating actions where true rollback isn't possible.
None of this makes the loop converge. It makes non-convergence cheap. That is a legitimate goal when you can't prove a loop will terminate: bound the damage per iteration instead.
What the research shows, and what it doesn't
I'll separate evidence from interpretation, since mixing them is how this topic gets overstated.
Foundational work on how loops operate. ReAct (Yao et al., 2022) established interleaved reasoning and acting. Reflexion (Shinn et al., 2023) showed feedback stored in memory can improve later attempts. AgentBench (Liu et al., 2023) evaluated LLMs as agents across multiple environments and reported long-horizon reasoning, decision-making, and instruction following as persistent difficulties. SWE-bench (Jimenez et al., 2023) tests agents on real software issues requiring multi-step interaction with a codebase.
Security work. AgentDojo and ASB, discussed above, are evaluations of tool-using agents under attack.
Standards. OWASP published the Top 10 for Agentic Applications 2026 on December 9, 2025, with risk identifiers ASI01 through ASI10. It provides a taxonomy and a shared vocabulary, covering goal hijack, tool misuse, identity and privilege abuse, supply chain, unexpected code execution, memory poisoning, inter-agent communication, cascading failures, human-agent trust exploitation, and rogue agents. The NIST AI Risk Management Framework's Generative AI Profile is a governance foundation for managing generative-AI risks across the lifecycle. It doesn't address infinite loops specifically, and I wouldn't cite it as evidence of them.
The most direct evidence. Hou et al. evaluated IAL-Scan on 6,549 LLM agent repositories. The tool reported 74 potential findings, and manual review confirmed 68 IAL failures across 47 projects, for a precision of 91.9%. The paper was posted to arXiv in early July 2026 and is a preprint, not a peer-reviewed publication.
The caveat that matters most: 91.9% is the precision of the detector, not the prevalence of loops. Of the 74 things the tool flagged, 68 were confirmed as real. It does not mean 91.9% of agents have loops. It also doesn't tell us how many real loops the tool missed, since precision says nothing about recall. The corpus is a set of public repositories, not a random sample of deployed systems, and static analysis can't observe runtime behavior. Secondary summaries of the paper mention false positives from dynamically resolved bounds and false negatives from custom termination logic and orchestration loops outside the supported frameworks. I'd treat those as plausible limitations rather than reported findings, since I'm relying on a summary.
What the paper does support, with appropriate scope: structurally unbounded agentic loops exist in real, public projects, they can be detected statically at reasonable precision, and the authors' framing of continuation control as the root issue is a coherent one. What it doesn't support is any claim about frequency in production or about any particular model.
The lineage runs in a recognizable order:
2022 ReAct Reason → Act → Observe
↓
2023 Reflexion Act → Feedback → Memory → Retry
↓
2023 AgentBench / SWE-bench Long-horizon agent evaluation
↓
2024 AgentDojo / ASB Security of tool-using agents
↓
2025 OWASP Agentic Top 10 Agentic security as its own domain
↓
2026 Infinite Agentic Loops Formal treatment of unbounded feedback paths
The IAL concern follows from the architecture's evolution. It isn't a new anxiety attached to an old system.
An experiment I would run
I haven't run this, and nothing below is a result. It's a protocol for what a rigorous test of the argument would look like.
The environment would be synthetic and harmless: a fake search tool, a fake database, a fake ticket system, and a synthetic state store. No real credentials, no production access, no real side effects. Tasks would be seeded so that some admit a solution and some are unsolvable, since the interesting behavior is what the agent does when it can't succeed.
| Config | Loop bound | Retry bound | Progress check | Repetition detection | What it isolates |
|---|---|---|---|---|---|
| Baseline | No | No | No | No | Raw failure behavior |
| A | Yes | No | No | No | Budget-only |
| B | Yes | Yes | No | No | Retry control |
| C | Yes | Yes | Yes | No | Progress-aware |
| D | Yes | Yes | Yes | Yes | Loop-aware |
| E | Yes | Yes | Yes | Yes (plus cost, timeout, authorization) | Full controller |
Measuring only "did it finish?" would miss most of what matters. The set I'd record:
| Metric | Why it matters |
|---|---|
| Task success | Utility |
| Average and maximum iterations | Typical vs. worst-case behavior |
| Tool calls per task | Capability consumption |
| Retry rate | Recovery behavior |
| Repeated-action ratio | Loop tendency |
| Context growth | State explosion |
| Token usage and cost per run | Economic impact |
| Time to stop | How fast containment engages |
| Unauthorized actions | Security |
| Side effects per run | Operational risk |
| False stops | Over-aggressive defenses |
| Recovery rate | Resilience |
The tradeoff I'd expect, as a hypothesis rather than a prediction of results, is that a system becomes safer by stopping more often, and that too much stopping destroys usefulness. The experiment's value would be locating where on that curve each configuration sits.
Who controls continuation
Return to the opening question: what guarantees the next transition is the last one?
For an unbounded system, nothing does. The model might emit a final answer, or it might not, and the surrounding code is built to follow whichever happens.
The problem is not that AI can loop. Humans deliberately build loops into intelligent systems, because loops are what make agents useful. The deeper problem is who controls continuation.
My answer is architectural:
Model proposes
↓
Runtime verifies
↓
Policy authorizes
↓
Tool executes
↓
State changes
↓
Progress is measured
↓
Runtime decides: STOP / CONTINUE
An agent should never be trusted to define its own stopping boundary when the system can independently enforce one.
The most important feature of an autonomous system may not be what it can do. It may be the part that makes it stop.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.







