Est.

Organizing Context in a Multi-Agent Harness

How to build a shared context layer that keeps multiple agents aligned.

Features Editor · · 10 min read
Cover illustration for “Organizing Context in a Multi-Agent Harness”
Agent Runtime Internals · October 3, 2026 · 10 min read · 2,286 words

Context management in a multi-agent system is an operations problem. A team of agents fails when two of them hold different, confident, internally consistent versions of the same fact, and nobody in the system is positioned to notice. Running agents reliably at scale, on a schedule, across sessions, means building a layer that governs which version of a fact is official, where it's visible, and when it expires, before any agent starts working.

Why multi-agent context management differs from single-agent memory

A single agent that runs out of context fails in a recognizable way: it asks a vague question, gives an incomplete answer, or hallucinates a fact it should have retrieved. That failure is visible because there's only one version of the truth in the room. The agent either had it or didn't.

Multi-agent systems don't fail that way. Another might use a slightly different definition, one that was approved six months ago and never retired. Neither agent is missing information. Both are using information that happens to conflict.

That distinction separates two things people tend to treat as one. Passing work forward is not the same as agreeing on what the work means.

Once that split is visible, the implication is hard to avoid. Context in a multi-agent harness is a shared operating layer that every agent reads from and writes into, and like any shared layer with multiple writers, it has to be engineered: versioned, scoped, and governed, the same way a team would treat a production database rather than a scratch file.

The five failure modes that collapse multi-agent coordination in production

Four of these failure modes are structurally impossible in a single-agent setup, because a single agent holds the only copy of context there is. They become possible the moment a second agent enters the picture, and they compound with every agent added after that.

Simultaneous-write conflicts round out the list. A GitHub issue filed against Letta proposes policy enforcement for read and write access along with audit logging, which is a real step toward memory governance, but it stops short of addressing what happens when two writes land at once.

A revenue workflow shows what it looks like when all five modes hit at the same time. Each one does its individual job correctly, by its own internal measure. The final report reads clean, specific, and confident. It's also wrong, and nothing in the pipeline was built to catch that, because each agent only ever verified its own step.

What a governed context layer must contain and decide

Fixing this starts with being precise about what a shared context layer actually is. It isn't a vector database. It isn't a prompt file sitting next to the model. It isn't a memory tool bolted onto an agent framework. All three of those can play a supporting role, but none of them decides which definition is certified, which version of a context object an agent used for a given answer, or which policy currently governs a given workflow. Those are governance decisions, and something in the system has to make them explicitly.

The artifact that encodes those decisions is often called a Context Repo: a bounded, versioned package containing definitions, policies, metadata, quality signals, and evaluation traces for a given domain. Treat it the way engineering teams treat a release artifact. It gets reviewed before it ships, tested before it's trusted, promoted when it's ready, and rolled back when it's not. Agents don't write freely into it the way they'd write into a shared scratchpad.

Validating what belongs inside that repo benefits from a clear checklist. Vera Vishnyakova, at HSE University, proposed five context quality criteria: relevance, sufficiency, isolation, economy, and provenance. Economy keeps the whole thing from becoming an unmanageable sprawl of context nobody actually uses.

What comes out of applying that discipline is a layered structure with four distinct kinds of content, each with its own lifespan and its own rules. A governed knowledge layer holds the business context, the semantic models, and the policies that don't change task to task.

Across all four layers, you want the smallest certified bundle that lets each agent do its job safely. Context management for multi-agent systems is fundamentally an operations problem: the system has to decide which version is official, which slice belongs at which step, and when each piece expires. UFO, built as an agent operating system with memory as a first-class primitive, centralizes that governance through a shared context layer agents coordinate around rather than pass between, treating persistence and proactivity as infrastructure decisions instead of something each team has to reinvent from scratch.

Four coordination architectures and the trade-offs that make the choice non-trivial

Governing the content of context is one problem. Getting the right slice of it to the right agent at the right moment is a separate one, and it's where coordination architecture comes in. Picking the wrong pattern for a given environment doesn't just fail to help, it amplifies the failure modes already described.

A centralized context broker routes every read and every write through one service. The cost is a single point of failure and a latency bottleneck, which rules it out for anything running at high concurrency. It fits regulated, low-volume workflows where auditability matters more than raw throughput, since every read and write routes through the single broker.

A hierarchical model splits context into a parent layer that holds the global picture and child layers that hold local detail. It fits regulated industries well, and it happens to be the dominant deployment pattern in 2026: a single orchestrator holding the full conversation context while spawning short-lived, isolated subagents that return only compressed summaries once they finish.

A distributed context mesh propagates changes through pub/sub events across agents. Hand-rolled pub/sub invites the exact context drift this whole discipline exists to prevent.

Peer collaboration with supervisor mediation has agents working concurrently over a shared bus, with a supervisor involved but not holding the only copy of context. It fits heterogeneous agent teams where no single orchestrator can reasonably hold the full domain, and it also produces the hardest simultaneous-write and conflict-resolution problems of the four patterns.

None of these four is a default correct answer. The choice has to match the environment, not the other way around, and each failure mode from earlier in this piece becomes easier or harder to trigger depending on which pattern a team picks.

MCP and cache-efficient context management for delivery and cost

Picking a coordination pattern settles the architecture, but it doesn't settle how context actually gets delivered to an agent at runtime, or what that delivery costs in inference spend, and both of those have become first-order engineering constraints rather than afterthoughts.

MCP, the Model Context Protocol, has gone through a real architectural split, and you need to understand it on its own terms. Durable execution answers a different one: how does an agent keep working reliably over time, across restarts and interruptions. Those two concerns used to get conflated inside individual frameworks. Separating them at the protocol level means a team can swap out an execution layer without touching how agents talk to tools, and vice versa.

Three MCP primitives carry the state-management work. Prompts package reusable templates that encode procedural memory, so an organization can distribute a standard incident-response or code-review template to every agent running against its systems, uniformly, instead of letting each team write its own version.

Cost comes in through a trade-off that's easy to miss until it bites you. TokenPilot, a paper posted to arXiv in June 2026 and revised in August of that year, lays out the problem directly: trimming a context window through text pruning or dynamic memory eviction reduces token count, but mutating the sequence changes the prompt's layout. That breaks the prefix match a KV-cache depends on, which invalidates the cache and erases the savings the trimming was supposed to produce. Sparsity and cache continuity pull against each other.

TokenPilot's answer runs on two levels. Ingestion-Aware Compaction works globally, stabilizing the prompt's prefix and filtering out noise at the point context enters the system. Lifecycle-Aware Eviction works locally, watching how much residual value a given context segment still holds and offloading it only once its relevance to the current task has actually expired. The point extends past this one paper: compression applied without accounting for prefix stability destroys the exact cache reuse that made compression worth doing. Any policy for shrinking context has to be built with cache layout in mind from the start, not layered on after the fact.

Durable state, memory governance, and the visibility trap that misleads engineering teams

The most dangerous mistake in this entire discipline is assuming that because a team can see what its agents are doing, it has control over what they do.

Agents working in observable channels, logging actions, posting to shared threads, surfacing outputs for review, create a natural but false sense that oversight is happening in real time. Engineers watch the activity and sign off on final outputs, and that feels like control. The Coalition for Secure AI has flagged exactly this pattern as a security risk, because an agent can take a consequential, hard-to-reverse action well before anyone reviews it. Watching a log after the fact is not the same as having a gate before the fact. Confusing the two is how irreversible actions slip through systems that looked, on paper, well-supervised.

Durable state is the infrastructure fix for part of that problem, and it has to be built in, not patched on. A workflow that stores pending runs in server process memory breaks the moment a load balancer routes a follow-up request to a different pod than the one that handled the original request. The run gets orphaned, with no record anywhere of what it was waiting on. Modern agent frameworks handle this with a pause-and-resume model: state gets serialized and stored at the exact point a human review is needed, and the same run picks back up once a decision comes in, regardless of which process or machine handles the resumption. That only works if the underlying infrastructure treats state as a durable, first-class artifact rather than something living in memory until the process restarts.

Memory governance covers the rest of the problem, and it belongs in the same conversation as compliance. Durable memory, the approved learning that persists across sessions, needs review cycles, privacy controls, and provenance tracking in any enterprise setting, the same as any other asset with regulatory exposure. Local MCP gateways have started to matter here because they let an organization host its memory layer on its own infrastructure while still querying cloud-based language models, so the most sensitive institutional knowledge stays inside its own security perimeter. High-privacy teams increasingly treat memory as a governed data asset subject to the same enterprise security protocols as any other sensitive store.

Each failure mode described earlier, the simultaneous-write conflicts, the silos, the staleness, the missing provenance, the silent divergence, traces back to treating context as something passed around between agents rather than something a governed operating layer actively manages. Systems built from the ground up as agent operating systems, UFO among them, bake transaction semantics, audit trails, and freshness guarantees into the coordination substrate itself, so these failures become structurally harder to trigger rather than problems each team has to solve on its own after the fact.

A working multi-agent context architecture in a deployed system

If an agent can't carry its context across sessions, or act against a durable state machine, you can't really call it autonomous. It's an expensive function call that starts over from zero every time it runs, no matter how sophisticated its reasoning looks in a single pass.

A deployed system that avoids that trap has a few concrete properties, and they tend to appear together rather than in isolation. Context gets scoped per agent role, so a given agent only ever gets the certified slice of a Context Repo it needs for what it's doing, not the entire domain's worth of definitions and policy. State persists through a durable execution layer that survives process restarts, pod rescheduling, and human-in-the-loop pauses, with runs resuming from serialized state rather than getting silently dropped. Delivery runs through MCP primitives, Resources for typed state, Roots for scoped file and path access, Prompts for shared procedural templates, so agents built on different frameworks can still interoperate without custom integration work for every pair. Compression and caching get handled together, with ingestion-time stabilization and lifecycle-aware eviction working in tandem so that shrinking a context window doesn't quietly wreck the cache reuse that made the system affordable to run at scale. Governance is built into the Context Repo itself: versioning, access rules, and audit trails live there, not bolted on as a separate compliance exercise after deployment.

The Context Repo pattern, a bounded, versioned, reviewable package of definitions, policies, and access rules, is the artifact that turns governance from an aspiration into something enforceable. Tool-connected environments that centralize cross-agent workflow execution and memory through a shared workspace model, UFO among them, naturally enforce this kind of structure: agents reach context through the platform's coordination layer instead of each fetching its own copy independently, which makes versioning, promotion, and rollback a property of the platform itself rather than a manual discipline a team has to maintain by hand. For anyone planning to run agents on a schedule, over long horizons, at real scale, the actual design question to ask of any platform under consideration is not whether it can call a model, but whether it treats context as a governed layer the system owns, or as a file each agent carries around and hopes stays current.

Sources

  1. Context Engineering: From Prompts to Corporate Multi-Agent Architecture
  2. TokenPilot: Cache-Efficient Context Management for LLM Agents
  3. The 2026-07-28 Specification

More in Agent Runtime Internals