Est.

Scheduling Autonomous Agents on Shared Compute

GPU schedulers must treat multi-step agent workflows as single units, not isolated requests.

Correspondent · · 10 min read
Cover illustration for “Scheduling Autonomous Agents on Shared Compute”
Agent Runtime Internals · October 1, 2026 · 10 min read · 2,204 words

A GPU scheduler looking at an agent sees a pile of separate requests: one prompt in, one completion out, over and over. What it's actually looking at is a single task that happens to be broken into pieces, each piece leaning on everything that came before it. That gap between what the scheduler sees and what's actually running is the whole problem.

Current GPU schedulers treat each LLM inference call as an independent, stateless unit, and that unit is the wrong one for agent work. An agent runs anywhere from tens to hundreds of chained LLM calls to finish a single task, and between each of those calls, schedulers throw away gigabytes of intermediate state, rebuilding context from scratch at every step. That habit inflates end-to-end latency by a multiple that static, single-call serving never has to deal with. Zhang et al., 2026 describe this as a deeper shift than a performance tweak: chatbot-era systems were built for fast, self-contained responses, while agents now hold persistent workspaces, reusable skills, and long trajectories of state, action, and observation stretched across an entire task. Systems designed for the first kind of cognition don't fit the second.

A natural question follows: if each individual call ran faster, would the overhead wash out? But the penalty here compounds across the chain of calls, not within any single one of them, so shaving milliseconds off inference doesn't touch the cost of reconstructing state at every step. The latency multiplier scales with how long the task is, not with how fast the model responds. Schedulers are optimizing the wrong variable, and that mistake produces everything else in this piece.

Structural differences between agent workloads and ordinary LLM serving

Agents are compound workloads, and three properties of that compound structure break the assumptions request-level schedulers depend on: state that keeps growing, branching that can't be predicted in advance, and tool calls that interleave unpredictably with everything else.

Take growing state first. Every multi-turn agent conversation grows its KV cache as it goes, and a 10-turn conversation with a large model consumes meaningfully more VRAM than a single-turn query ever would. Infrastructure sized around a model's parameter count alone will run short the moment real agent sessions start piling up context. A three-agent system, say an orchestrator paired with a coder and a reviewer, needs a substantial pool of combined VRAM as a baseline, and that's before any KV cache growth even enters the picture. Size for the model, and the model fits. Size for the agent, and the math changes.

Then there's the orchestration layer itself, and this is where a lot of capacity planning goes wrong. Orchestration work, scheduling sub-tasks, routing tool calls, passing state between agents, is CPU-intensive work, and it demands a CPU-to-GPU core ratio far richer than a typical inference cluster carries, somewhere around 1:1 to 1.4:1. CPU processing can account for the majority of total latency in agentic workloads, an analysis cited by GPU Mart found. For teams coming from static LLM serving, that's the single most surprising figure in the whole migration. Dynamic branching adds to the burden here too, since the orchestration layer is CPU-heavy, not just GPU-heavy.

Tool-call interleaving closes the loop. Every tool call an agent makes has to be routed somewhere: immediate GPU execution, a queued slot on the GPU, or offloaded to CPU entirely, all under a shifting VRAM budget. Two runtime factors, GPU utilization contention and VRAM capacity contention, are what push real, measured latency away from whatever static profile was planned in advance. Tool calls occur while the agent is already running, so no schedule computed before the task started can know what resource demand is actually coming. AgenticOS names the root cause: processes, threads, files, sockets, the whole vocabulary of traditional operating systems, was built for workloads that don't change shape mid-flight. Scheduling primitives borrowed from that vocabulary inherit the same blind spot.

What program-level scheduling requires the runtime to track

If the inference call is the wrong unit, what's the right one? Recent research answers directly: the program, the whole agent workflow from start to finish, is what the scheduler needs to reason about, not any individual call inside it. That single move, treating the workflow as the schedulable unit instead of the request, is the pivot the rest of this argument depends on.

Making that pivot real requires the runtime to track four things it currently doesn't. The agent needs a stable identity that persists from its first step to its last, so the scheduler can apply resource limits, priority, and isolation to the whole workflow instead of patching each fragment separately. The runtime needs to hold accumulated state, meaning intermediate outputs, KV cache positions, tool call results, and approval receipts, as data the scheduler can actually address, rather than something rebuilt from zero on every call. It needs a forward-looking resource envelope for the workflow as a whole, one that updates as tool calls reveal real demand instead of estimating each call in isolation. And it needs lifecycle operations, spawn, suspend, resume, terminate, that mean something at the level of a program. None of those four operations mean anything applied to a single inference call.

Two different research efforts arrive at this from opposite directions, and they converge anyway. AIOS proposes an LLM-agent operating system that isolates memory, storage, context, scheduling, and access control into a kernel that manages concurrent agents directly, making the agent workflow the kernel's managed unit, analogous to a process. Quine takes a different route to the same place: it realizes agents as native POSIX processes, where identity is just the PID, lifecycle is fork, exec, and exit, and state splits cleanly across process memory (ephemeral), environment variables (scoped), and the filesystem (persistent). That approach inherits kernel-level isolation for free, without inventing a new abstraction layer on top of the one operating systems already have.

The Chatbot-to-Digital-Colleague framework gives a sense of what's actually at stake in getting this right. Its "Workspace + Skill" paradigm turns episodic tool use into something closer to how a colleague works: state persists, procedures get reused, tasks actually close out, and past experience carries forward into the next task. None of that is visible to a scheduler that only ever sees one request at a time.

How memory architecture determines whether program-level scheduling is possible

Diagram: Three Memory Tiers a Program-Level Scheduler Must Track. Visualizes: Show three distinct memory tiers that a program-level scheduler must reason about separately, arranged as a vertical stack or layered diagram moving from volatile to…

Program-level scheduling sounds achievable on paper, right up until an engineer asks where the program's state actually lives between calls. Large language models carry no built-in memory from one call to the next. Any production agent that spans more than a single request needs a deliberate memory layer built underneath it; without one, the agent forgets everything the instant its context window resets. Without that layer, program-level scheduling has nothing durable to schedule around, no matter how well-designed the scheduler itself is.

The failure mode is recognizable to anyone who has run an agent past the prototype stage. Context windows balloon because the same information gets reconstructed over and over. A retry has no memory of what happened right before it failed, so it starts over blind. Approvals and other intermediate decisions get buried in chat history instead of living somewhere queryable. Multiple copies of what should be one source of truth drift apart until nobody can say which one is current. The fix that 2026 guidance points to is a clean division of labor: keep only the current turn's working memory inside the model's context window, and move everything durable, retries, approvals, progress tracking, into systems actually built to store and query that kind of state.

That division maps onto three distinct memory tiers a program-level scheduler has to reason about separately. In-context working memory holds the current reasoning chain, is volatile, and lives entirely inside the model's context window. External short-term state holds current task progress, tool results, and intermediate outputs, kept outside the model where it can survive a context reset. Long-term persistent memory holds prior task outcomes, learned preferences, and past decisions the agent needs to recall across sessions that may be days or weeks apart. A scheduler that can't distinguish between these three tiers has no way to decide what to preserve, what to discard, and what to hand back to the model on the next call.

Bayer's use of Cognee to run scientific research workflows shows what this looks like at production scale. Its agents need to recall prior experimental findings, connect new data to existing hypotheses, and maintain provenance across research pipelines that run far longer than any single session. A scheduler treating each step as independent simply cannot serve that workflow, because the entire value of the agent lives in the context it accumulates across sessions, not in any single call. Deloitte estimates that by 2027, roughly half of companies using generative AI will be running agentic AI pilots or proofs of concept, up from a quarter in 2025. The agents built during that wave will need memory architecture designed for production from the start, rather than the prototype-grade shortcuts that work fine for a demo and fall apart under real load.

Shared compute and amplified scheduling problems across concurrent agents

Everything so far has been about a single agent running on infrastructure sized for it. Putting several agents on the same cluster turns the same structural problems from isolated incidents into compounding ones. Request-level scheduling, already a poor fit for one agent, produces correlated contention spikes across several agents that no single one of them can see coming or absorb on its own.

The two contention factors named earlier, GPU utilization contention and VRAM capacity contention, don't just persist under shared load, they get sharper, because now multiple agent programs are drawing on the same finite pool at once. KV cache growth is the clearest example of why this isn't a simple additive problem. Each agent's cache eats into a shared VRAM pool, so one long-running agent can quietly starve every other agent on the same cluster mid-execution. A scheduler working at request granularity has no way to see this coming, because it never had a view of any agent's program-level VRAM trajectory in the first place.

So what actually resolves contention that no single agent can see? MAS-DecStream offers one answer worth taking seriously. The research shows that static, single-stage scheduling decisions made under partial system visibility lead directly to resource overcommitment and QoS violations. Its proposed fix is LLM-assisted, multi-round Contract Net Protocol negotiation, where edge-cluster agents iteratively refine their offloading proposals using local observations, predicted resource states, and qualitative context about what's happening elsewhere in the system. Hard resource and QoS constraints stay deterministic throughout, but the qualitative judgment calls, which task matters more right now, which cluster is trending toward saturation, are exactly where negotiation beats a fixed rule set. That's program-level scheduling extended across agents rather than confined to one: each agent exposes its own resource intent and timeline, and negotiation resolves conflicts before they turn into latency violations, rather than after.

Work by Sochat and Milroy at Lawrence Livermore National Laboratory grounds the same idea in a real production context. When dispatch agents receive jobs described in natural language with rich metadata, instead of rigid structured scripts, they can match each workload to the resource actually suited to it, rather than defaulting to whatever happens to be free first. That single change eliminated architecture mismatches and improved performance across most of the applications tested. With descriptive metadata in place, successful job execution reached 87% in their experiment, compared with under half without it.

The gap between theory and runtime reality in production deployments

Production teams are solving these problems themselves, without waiting for a purpose-built agent operating system to arrive. They're solving pieces of it themselves, often without naming what they've built, and the shape of their workarounds tends to match, almost exactly, the runtime properties a purpose-built agent OS would provide natively.

incident.io's AI SRE is a clear case. It embeds directly into Slack workflows, and when an alert fires, it automates up to 80% of incident response by identifying the likely change behind the incident, suggesting next steps drawn from past incidents, and pulling metrics and logs from Datadog or Grafana straight into Slack. That last part cuts down, though it doesn't eliminate, the need to context-switch into those separate tools. Look closely at what that system actually requires to work: it needs to recognize the same incident across multiple steps, hold state about what's already been tried, and pull in historical context from past incidents to inform what happens next. That's identity, accumulated state, and long-term memory, the same three things program-level scheduling names as requirements, just built by hand rather than provided by a kernel underneath.

That's the honest state of the field right now. The theory of program-level scheduling, treating the whole agent workflow as the unit that gets identity, memory, and resource intent, is sound, and the production examples across incident response, scientific research, and edge computing all confirm the same underlying shape. What's missing in most deployments is the infrastructure to apply that insight by default, instead of rebuilding it, one integration at a time, inside whatever workflow tool happens to be closest at hand.

Sources

  1. Descriptive Dispatch of Computational Work
  2. Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
  3. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
  4. How to Build GPU Infrastructure for AI Agents: The 2026 Compute Playbook
  5. AgenticOS 2026: Operating Systems Design for AI Agents
  6. Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents

More in Agent Runtime Internals