Est.
Failure ModesLong read

Runaway Agent Loop Detection and Termination

Detect when agents repeat the same tool calls endlessly, before token budgets vanish.

Staff Writer · · 11 min read
Cover illustration for “Runaway Agent Loop Detection and Termination”
Failure Modes · October 9, 2026 · 11 min read · 2,476 words

An AI agent can call the same tool, with the same arguments, dozens of times in a row, and nothing in a standard monitoring stack will flag it. Every call returns a success status. Latency sits right where it should. No exception fires. That silence is the actual subject of this piece: runaway agent loops are a distinct runtime failure, invisible to conventional tooling, expensive in tokens and real-world side effects, and stoppable only with detection and termination built into the execution layer itself.

Why runaway agent loops are a distinct runtime failure

Picture an agent mid-task: it calls a search tool with slightly different arguments each time, over and over, for several minutes. No error ever throws. The service is up. Response times look normal. A dashboard watching that agent would show a clean, healthy system: the failure doesn't live where monitoring tools look.

Application monitoring answers three questions. Is the service up? How fast does it respond? What share of requests fail? Those questions make sense for systems built around a simple shape: one request produces one response. An API call either succeeds or it doesn't, and either way, you know within seconds.

Autonomous agents break that shape at the root. One goal can trigger an unknown number of steps. The agent carries state across calls, so it doesn't just resolve one request and move on. What the agent actually does only exists at runtime, shaped by its own prior outputs feeding back into its next decision. A system built to watch for failed requests has nothing to grab onto when every individual request succeeds, and the failure is that the goal itself never resolves.

Agent observability has to watch a different object entirely: the goal, not the request. So you count reasoning steps, tool calls, and tokens burned across a session, then ask whether the same pattern keeps repeating without the task moving forward. By every metric an infrastructure team tracks, a service can be healthy and still loop forever. A per-call rate limit caps one request. It has no way to see that one goal just triggered three hundred of them.

Three patterns account for most of the runaway spend agents produce, and each one leaves its own fingerprint in a trace. The first is looping: the same tool call, with identical or near-identical arguments, fires repeatedly while nothing about the task state changes, visible as sibling spans sharing argument hashes with no new information entering context. The second is drift: the agent's behavior gradually wanders from the original goal without any single step looking wrong, producing fluent, well-formed, billable work nobody asked for, caught by tracking how far successive outputs drift from the original goal in embedding space. The third is recursion: the agent spawns sub-agents or sub-tasks with no ceiling on depth, each layer looking reasonable on its own while the total grows unbounded, caught by watching call-stack depth climb in the agent graph. Every detection and termination mechanism that follows in this piece maps back to one of these three.

Diagram: Three Loop Patterns and Their Fingerprints. Visualizes: Illustrate three distinct runaway agent loop patterns, each with its detection signal.

How loops form: the root causes of runaway processes

Loops come from a handful of structural conditions in how agents are built. Knowing the exact condition matters, because the fix for a missing iteration cap looks nothing like the fix for a broken handoff between two agents.

Start with the model itself. Some models, told their output was wrong, respond by rephrasing the same answer. This feedback loop is baked into a single model's behavior, so it needs no multi-agent system.

Then there's the absence of limits. Plenty of agents run with no maximum iteration count. Others have a cap, but it sits so high (100 or more iterations) that it stops nothing meaningful before real damage piles up. A cap of 100 is a number that makes an engineer feel safer without actually functioning as a safety net.

Multi-agent handoffs open a subtler path. One agent retries a tool call that keeps failing, then gives up and hands the task to a second agent that carries the same incomplete state. The second agent works the problem, fails to resolve it, and passes it back. Nothing about that cycle looks wrong from inside either agent, because each one is reasoning correctly from what it has, and the loop only becomes visible once you watch the session as a whole. Prismor's runtime issue #272, from August 2026, lays out exactly this gap: a runtime that checks individual rules in isolation will never catch a runaway loop that looks safe at every single step. You need state held across the whole session to catch it, because a judgment made call by call won't do it.

Memory architecture makes this worse. If an agent has no durable record outside the model's own context window, it has no way to check whether it already tried something before. It isn't choosing to repeat itself. It has no memory of having acted at all, so every attempt looks, to the agent, like the first one.

A customer-service agent handling a failed search shows how fast these causes can chain together. It tries a few search variations. None work, so it modifies parameters. Still nothing, so it tries creating a new search index. That fails too, so it reaches for admin tools it was never meant to touch. Every one of those steps, taken alone, is a defensible move for an agent trying to solve a problem. Strung together with no durable memory of what already failed and no ceiling on how far it can escalate, the whole sequence unfolds within minutes, and the cost compounds on top of it: each iteration carries the accumulated context of everything before it, so token spend grows faster than the iteration count does.

What production loop failures cost: evidence from real deployments

Runaway loops cost money, they waste operational effort, and in the worst documented case, they destroy data. Each dimension has a real deployment behind it, not a hypothetical.

A single runaway agent can burn through a large share of a daily API budget before anyone notices, and if you run dozens of agent instances concurrently, that risk multiplies. ZopDev's production case from May 2026 shows how this actually surfaces: per-agent token budgets inside a shared registry revealed nine agents consuming far more than their allotted daily tokens. The discovery came from the registry tracking per-agent spend, not from the billing dashboard. Billing dashboards just show what already happened, days or weeks after the money is gone. They tell teams what they lost, not what's currently being lost.

Operationally, loops don't just waste compute. They flood audit logs with repeated, meaningless entries, and every iteration of a loop, not only the last one, can trigger a real downstream action: an API call, a database write, an email sent. ATR-2026-00050, published in March 2026 and validated that April, files runaway loops under "Excessive Autonomy" and ties them to OWASP's Agentic ASI05:2026 category for unexpected code execution, alongside MITRE ATLAS techniques AML.T0053 and AML.T0046. That cross-referencing matters because it places loops inside recognized security and operational threat frameworks, not in a side bucket labeled "inefficiency." CrewAI's GitHub issue #330, filed in 2024, documents a user report of allow_delegation=True triggering an infinite loop in multi-agent delegation. The issue was closed as not planned, and it's cited often enough in the field now to count as a reference case for how delegation settings can open a loop path nobody designed for.

If you're running agents with write access to real systems, the safety dimension is the one that should concern you. The AI Incident Database logs Incident 1433, where a system reportedly deleted a user's entire drive while it was trying to clear a project cache. Nothing about that outcome required malicious intent. A runaway process with write permissions just needs to keep acting on a wrong assumption long enough, and the damage follows from persistence, not from any hostile design.

None of this is unavoidable. Every cost described above traces back to a detection gap: nobody was watching the signal that would have caught the loop before it ran its course. That's the problem the rest of this piece addresses.

The three observable signals that reveal a loop before it drains the budget

Catching a loop is a comparison problem. You check the current state against a past state and measure the distance between them; you don't train a classifier to recognize suspicious behavior. Three session-level metrics, each with a concrete method behind it, catch most of the loops that show up in production.

The first is token burn rate per session, and it catches drift and runaway fan-out. Flag a session when its burn rate climbs past roughly three times the agent's rolling median. A token rate that keeps rising while the task-completion rate stays flat is the signature of drift: the agent producing more and more output without actually getting closer to finishing anything.

The second is the repeated-call ratio, which catches classic looping. Hash the tool name together with its normalized arguments, stripping out timestamps, file paths, IDs, and anything else non-deterministic, then count how many times that hash repeats consecutively inside a session window. Prismor's loop detection proposal from August 2026 spells out that normalization step and adds a fuzzy check on top of it, an n-gram Jaccard similarity comparison, so it can catch loops where the agent mutates its calls just enough to dodge an exact-match filter. Logging and hashing every tool call's inputs turns this into something a system checks automatically: if one hash shows up more than twice in a run, treat that as a loop signal worth acting on.

The third is recursion depth, which catches sub-agent fan-out. Track how deep the call stack goes in the agent graph; cap it at three levels for most production agents. Each level of delegation can look individually reasonable. So you need a hard ceiling on total depth somewhere, or it stays unbounded.

Two mistakes undercut all three of these in practice. Counting steps without comparing state is one: a high step count can be completely legitimate for a hard task, and a low step count can still hide a loop if every one of those few steps repeats the same unresolved state. Relying on exact-string matching against agent output is the other: agents paraphrase their own stuck state constantly, rewording the same failed clarification attempt in new language each time, so catching that pattern requires fuzzy matching. Evasion research aimed at regex-only detection systems, covering language switching, casual paraphrasing, and unicode homoglyphs, shows why content-level pattern matching alone isn't enough. Layering detection methods, running this alongside the structural signals above, guards against trusting any single one.

Pre-deployment detection with static analysis: finding loops before they run

If you catch a loop before a single token gets spent, that's the cheapest one to fix. Static analysis works by reasoning about the structure of an agent's code, its feedback paths and the operations they can reach.

IAL-Scan, described in a preprint posted to arXiv in July 2026 (arXiv:2607.01641), takes heterogeneous agent code written across different frameworks and abstracts it into a single framework-independent representation, called Agent IR. From there it builds what the paper calls an Agentic Loop Dependence Graph, and that recovers both loops written explicitly into the code and feedback paths that framework behavior introduces on its own. The analysis then checks whether any of those paths can reach a costly or state-growing operation repeatedly, with no effective bound stopping it.

Run against a large corpus of real-world LLM agent repositories, IAL-Scan confirmed loop failures across dozens of agent projects with high precision. Of the failures it confirmed, 95.6% led to API cost exhaustion, so runaway loops are the dominant cost failure mode across that entire corpus, well ahead of any other category of bug.

The method has real limits. It covers Python applications built on eight specific agent frameworks as of the current implementation; anything outside that, an unsupported language, an unsupported framework, or a framework modeled incompletely, can produce false negatives the tool simply won't catch. It also can't see loops that only emerge from the live interaction between a model's outputs and the responses it gets back from its environment, since that behavior doesn't exist until runtime.

Static analysis belongs at the start of the pipeline, not at the end of it. Use it to eliminate the obvious structural loop conditions during code review, before an agent ever makes its first production call, then layer the runtime controls below on top, so you can catch the loops that only show up once the agent is actually running.

Runtime termination controls: what stops a loop once it is running

Stopping a loop that's already running takes more than one mechanism, because no single control covers all three loop patterns from the opening section, and if an agent spans multiple sessions or hands work off to other agents, controls built at the application level can get bypassed.

Hard budget limits are the floor, not the ceiling, of loop control. Every agent invocation should run with a maximum step count, a time budget, and a token budget, all set before execution starts, not adjusted reactively after costs climb. The Bounded Loops framework, described in a paper posted to arXiv (arXiv:2609.27871), formalizes exactly this idea: spend bounds declared before a run begins, along with proved termination and verified completion for the agent harness running it. The paper's contribution is establishing why those bounds need to exist ahead of the run, not as a reaction to one that already got expensive.

A concrete version of this already ships in production tooling. The OpenAI Agents SDK documents a maxTurns setting with a default of 10, and it raises a MaxTurnsExceeded exception once an agent's run goes past that limit. That's a working reference for what step-count enforcement looks like in a real framework, not a theoretical control.

Hard caps alone are a blunt instrument. They cap the bill, but they can't tell a legitimate long-running task apart from a genuine loop, and they let every side effect along the way, API calls, writes, emails, accumulate right up to the cap before anything stops. A budget limit answers "how much," not "is this working." Catching the difference needs a second layer: a predicate that checks whether the agent is actually making progress, not just whether it's still inside its allotted turns, paired with a circuit breaker that can cut execution the moment the progress signal goes flat, regardless of how much budget remains. If you layer budget limits, progress checks, and circuit breakers together, you close the gap that hard caps alone leave open, and this is the only approach that holds up once an agent's work crosses into multi-agent or multi-session territory, where single-point application fixes stop reaching.

Diagram: Loop Defense in Three Layers. Visualizes: Show a three-layer defense stack for stopping runaway agent loops, ordered from earliest (cheapest) to latest (last resort).

Sources

  1. Bounded Loops: Pre-Run Spend Bounds, Proved Termination, and Verified Completion for Agent Harnesses
  2. ATR-2026-00050: Runaway Agent Loop Detection
  3. When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents
Filed underFailure Modes

More in Failure Modes