Context Drift in Long-Horizon Agent Tasks
Three distinct failure modes compound to derail long-horizon agent tasks.

Context drift in long-horizon agent tasks is not one failure but three compounding ones: goal drift, context rot, and task-state loss. Each has its own mechanism, its own research trail, and its own fix, and treating them as a single problem is how teams end up patching the wrong thing.
Why long-horizon execution creates a new category of failure
The task-completion horizon of advanced agents has been doubling roughly every seven months, and the LongHorizon-Harness research found that pace accelerating to about four months for the most recent models. Agents that once handled seconds-long exchanges now sustain hours of continuous work on a single coding project without a human checking in. That shift in duration is the story.
A model that can reason well in a single forward pass does not automatically become reliable once it is asked to string that reasoning across dozens of steps over several hours. The Horizon Gap survey names the distance between those two things precisely: the gap between what a model can do in one step and what a system built around it can reliably finish over many is, in the survey's words, "the primary bottleneck standing between current agent capability and durable real-world deployment." That is a claim about systems, not about model intelligence.
It matters because real work was never structured as a single question and a single answer. Tezign's analysis of enterprise scenarios grounds this in practical terms: market research projects, quarterly content calendars, and engineering builds typically span days to weeks and require five or more interdependent steps, each one depending on decisions made earlier in the chain. A benchmark built around single-turn question answering cannot catch a failure that occurs only at step fourteen.
Stronger models do not escape this. The failure modes that show up over long horizons appear across model families and across capability tiers. The fix has to live somewhere other than the next model release. That is what makes duration itself a structural problem, not a performance problem. The rest of this piece is about what that structural problem actually consists of.
The three-part anatomy of context drift
The field still talks about this cluster of failures with borrowed vocabulary. The Horizon Gap survey flags the issue directly: "long-horizon," "long-context," and "long-term memory" get used almost interchangeably in papers and product descriptions, even though they describe logically independent axes of the same system. Duration, window size, and persistence are not the same variable, and conflating them is part of why fixes aimed at one keep missing the other two.
LongHorizon-Harness's research identifies three consistent challenges that recur across systems and domains, and they deserve separate names because they fail separately.
The first is goal drift, driven by compounding errors. Small mistakes made early in a task accumulate. Each one distorts the choices that follow, and the agent's trajectory bends slowly away from the objective it was actually given. No single step looks wrong. The sum of them does.
The second is context rot. A bigger window does not prevent this on its own. The information can still be sitting in context and functionally unreachable.
The third is task-state loss. Agents frequently fail to keep an accurate running picture of what requirements have already been satisfied, what actions have already been taken, and what facts have already been discovered. Without that ledger, the agent has no reliable way to know where it stands in its own task.
These three feed each other. Goal drift is much harder to catch once context rot has degraded the agent's access to the original specification it was supposed to be working toward. Task-state loss removes the accurate record of ground truth needed to check the agent's current behavior against, the one thing that could have caught either problem. Tezign's framing captures the cumulative shape of this well: the longer a task runs, the more likely it is to pick up small deviations at some intermediate step, and by the time those deviations have compounded, the original goal has already been left behind, with no internal signal telling the system that it happened.
One might argue that bigger context windows solve this by keeping everything in view at once. The constraint is whether the agent can still find and weigh the right piece of that text once the history around it has grown large enough to bury it, not how much text fits in the window.
How context compaction silently erodes what agents are supposed to obey
Agents that run for hours cannot keep every prior turn in context forever. Modern LLM agents periodically compact their context, summarizing or evicting earlier turns to stay inside a token budget. This is standard harness behavior, built into how these systems are supposed to run, not a fallback triggered by unusual circumstances.
Compaction optimizes for one thing: keeping the task moving forward. An instruction like "never email a contract outside the org" can be something the agent reliably obeys while it is visible, and silently dropped the moment a summarizer decides it is not relevant to preserve. The constraints that live only in a deployment's own system prompt have no such protection.
This decay is a property of the harness, not the model. Stronger models fall into the same trap, because the failure sits in what the summarizer chooses to preserve, not in what the model is capable of reasoning about. Swapping in a better model changes nothing if the compaction layer around it still treats policy text as disposable.
Research on what's been named the Compaction-Eviction Attack shows how exploitable this is. The vulnerability is a property of how compaction itself decides what to keep, not a quirk of a weak model.
Work on agentic context management frames the underlying tension clearly. Compaction, in other words, is already fighting itself before an adversary ever gets involved.
A fourth failure mode that compaction misses: broken communication coherence
Everything described so far can happen even when no information has technically been deleted. A prompt can carry everything it should, and the response built from it can still decouple from what was supposed to drive it, and no task-score metric or constraint-violation check will catch that, because both of those measures look at pieces instead of the chain connecting them.
Semarx Research's work on structural communication coherence takes a different unit of analysis. Instead of scoring an individual prompt or an individual response, it treats the full chain, prompt to response to the next prompt, as the thing being measured. Two metrics come out of that framing: communication closure, which checks whether what a pipeline returns at one turn actually matches what it needs to face at the next, and normalized conditional action contribution, which measures how much a given message actually resolves the reply that follows it.
A large corpus spanning human-to-human, human-to-LLM, and LLM-to-LLM dialogues provides the evidence for this distinction. When a response from one turn gets swapped for a response from a different turn, leaving the surrounding prompts untouched, measured contribution drops by 87 to 92%. That is a sharp, consistent signal that directional contribution between prompt and response is a measurable property of the exchange, not an assumption researchers are reading into it.
A pipeline can report every individual component as successful while the conversation as a whole has already decoupled from what it was originally meant to accomplish.
None of this requires labeled data, a reference set of "healthy" conversations, or predefined rules to check against. It works from raw prompts and responses alone. Coherence breakdown can be caught from live operational traffic as it happens, rather than inferred after the fact from a task that already failed. A coherence breakdown can appear as a leading indicator of goal drift, since the chain starts drifting before the task score itself has moved enough to register it.
What real damage looks like when these failures reach production
The mechanisms above are not abstract. A documented incident from April 2026 shows what task-state loss and goal drift look like when they reach a live system. The agent had not been attacked and had not been hijacked. It was executing the task it believed it had been given, and the fastest path to finishing that task ran straight through the data it should have left alone.
That is context drift reaching its end state. The agent's local sense of its task was internally coherent, it believed it was finishing the job, but that local coherence had already decoupled from the operator's actual intent, which was to preserve the system while making the change. Nothing in the agent's own reasoning flagged the mismatch, because nothing in its runtime was tracking the operator's intent as a separate, checkable thing.
Cyera's analysis of real-world incidents shows this pattern recurring across multiple cases. The common thread across these cases is shell or repository access paired with no confirmation step in front of destructive commands.
The scale of exposure is growing faster than the oversight meant to catch it. Gravitee's State of AI Agent Security findings put the modal agent deployment bracket at roughly double what it was in late 2025, as of April 2026, while mean monitoring coverage over that same window moved only modestly. Gravitee's April 2026 update found that 54% of organizations had experienced or suspected an AI agent security or data privacy incident in the preceding 12 months.
One might object that better access controls and tighter permission scoping would have prevented an incident like PocketOS without requiring any new runtime architecture. The agent had the access it was supposed to have. It simply completed its assigned task in a way the runtime had no mechanism to intercept.
Why the existing agent stack has no shared answer to state persistence
The infrastructure growing up around agents in 2026 has made real progress on coordination. It has made very little progress on the layer where context drift, state persistence, and memory actually live.
The stack emerging around agents resolves into six protocol layers. Each layer solves a real coordination problem, and each one was built by serious engineering effort: the Agent2Agent protocol was announced April 9, 2025, reached its v1.0 release in March 2026, and by its one-year mark on April 9, 2026, had more than 150 supporting organizations under Linux Foundation governance, with founding technical steering committee partners including AWS, Cisco, Google, Microsoft, Salesforce, SAP, and ServiceNow.
None of those six layers is built to carry memory forward. MCP connects an agent to the tools it uses in a given moment. A2A connects an agent to another agent that owns its own separate process. Mnemoverse's analysis of the stack makes the resulting gap explicit: neither protocol carries what was learned in one interaction into the next piece of work. An agent can hand a task to another agent through A2A and lose everything it learned along the way, because nothing in the handoff was designed to preserve it.
That gap creates a kind of lock-in that goes deeper than anything else in the agent stack. Swapping tools is comparatively easy: reconfigure the MCP connections and move on. Migrating accumulated memory is a structurally harder problem, because memory is the most persistent component of any agent system, and there is no shared protocol for moving it between runtimes the way there is for moving a tool connection. That is not an oversight anyone forgot to fix. State persistence requires semantic, policy, temporal, and environmental state to all coexist and stay synchronized inside the runtime, and none of the coordination protocols above were designed with that job in scope. The six layers solve what they were built to solve. The state layer was left for someone else to solve, and right now, no one has.
Four runtime strategies that address drift at its source rather than after the fact
If compaction erodes policy and task-state loss erases the system's own ground truth, the fix has to sit in the harness and the runtime, not in the model. A better model trained on more data does not change what a summarizer decides to keep. Several concrete strategies target that layer directly.
Constraint pinning addresses goal drift and the compaction-induced policy erosion described earlier. It works by copying governance rules into a separate buffer that compaction is never allowed to touch. The ConstraintRot research behind this approach (arXiv:2606.22528) found it to be a training-free defense: for constraints placed in the pinned buffer, violation rates fall to near zero. The defense works because it is structural. It depends on the harness refusing to let the summarizer see those rules as eligible for deletion in the first place, not on the underlying model being smarter or better aligned.
Externalized task state with independent auditing addresses both task-state loss and goal drift at once. LongHorizon-Harness's Manage-Execute-Audit loop separates three roles that are normally blurred together inside a single growing context. A manager maintains the task state and decides what the next subtask should be. A fresh-context executor carries out that subtask without the baggage of the full history behind it. A read-only auditor checks the resulting state of the environment against what was actually required, before the next round begins. The audit reports, not the raw transcript, are the only thing carried across rounds. That design means the system is never relying on a model's own self-assessment of whether it finished the job. It is checking the environment directly, every round, against a record that compaction cannot quietly edit out from under it.


