Agent Lifecycle Primitives and State Machines
Explicit state machines make agent behavior auditable and predictable.

Once an agent makes more than one model call, it has a lifecycle. An engineer can model that lifecycle on purpose, or let it happen by accident, scattered across prompt history and tool output where nobody can audit it.
Why an agent needs an explicit lifecycle model
A single-turn LLM interaction is simple. The model produces text, a human reads it, and the human decides what happens next. There's no state to track, no sequence to protect, nothing to manage between steps because there are no steps.
That changes the moment a second model call enters the picture. Adding a tool invocation, a branch, or a follow-up decision that depends on what happened earlier gives the system state it needs to carry forward. That state has to live somewhere.
When nobody builds an explicit home for that state, it doesn't disappear, and teams run into trouble. It leaks into the prompt history and the tool output, and that's where non-reproducible behavior starts creeping in: the agent behaves differently across runs of the identical task because the "state" was really just whatever text happened to accumulate in context.
NJ Raman's Perceive → Reason → Act → Observe loop gives the minimal formal shape of what's actually going on. An agent runs this loop across an extended stretch of time, and at every turn it has to retain, or reconstruct, enough context to act coherently with what came before. That's the definition of what makes something an agent instead of a chatbot with extra steps.
Break that loop down further and the runtime process of a production agent resolves into five distinct stages: initialization, input, memory, decision, and execution. Each one can fail on its own, independent of the others. A memory failure can appear several steps downstream as a bad decision instead of a memory error. It might look like a bad decision three steps downstream, which is exactly the kind of thing that makes implicit lifecycles so hard to debug.
The engineering response to this has been remarkably consistent. Independent teams building open-source runtimes in 2026, LangGraph, CrewAI, and others, converged on the same answer without coordinating: model the agent as a graph, where states and transitions are explicit objects instead of implicit side effects of a prompt. That convergence is what happens when enough engineers independently rebuild the same missing primitive and arrive at the same shape for it.
State machines versus prompt loops
A state machine draws a hard line between what the agent is doing right now and how it decides what to do next. That separation is the single structural property that makes an agent's behavior something you can predict, debug, and test, rather than something you observe and hope for.
Look at what a plain prompt loop asks a single model to be. It's the reasoner. It's also the router, deciding what happens next. It's the state store, holding whatever context survived the last few turns. And it's the error handler, deciding what to do when something breaks. In a prompt loop, the LLM is simultaneously the reasoner, the router, the state store, and the error handler, all of which collapse into the same forward pass, the same next-token prediction. One model, wearing four hats, with no way to inspect which hat it's wearing at any given moment.
A state machine peels those roles apart. The LLM becomes one decision-maker operating inside a structured graph. It reasons within a given state, but the transitions between states, the guards that block certain transitions, and the persistence of what happened get handled by the harness surrounding it. The model does the thinking. The harness does the bookkeeping.
Why does that separation matter enough to build an entire architecture around it? Because it means the harness can be audited independently of the model. If something breaks, an engineer can look at the graph and the transition history without needing to reverse-engineer what the model was "thinking" partway through a context window.
There's a concrete number behind this, and it's a big one. A practitioner report from Chier Hu found that swapping an identical model out of a generic scaffold and into a purpose-built harness produced swings of more than thirty points on the same tasks. The model never changed. The structure around it did. That's the whole argument for state machines in one data point: performance was sitting in the harness the entire time, not in the weights.
The same report documented something almost counterintuitive. Stripping out the majority of an agent's available tools raised its success rate and cut latency at the same time. A smaller, more constrained state space beat a bigger, more permissive one. Constraint, applied at the right layer, is itself a performance feature.
What does an explicit state machine actually hand an engineer once it's built? Enumerable states mean every situation the agent can be in can be listed out, which is the precondition for writing a real test suite or drawing an error boundary anywhere at all. Inspectable transitions mean the question "why did the agent move from A to B" gets answered by reading the graph, not by re-reading long stretches of prompt history looking for clues. Guards on transitions let a team enforce a rule like "get approval before this destructive action runs" as a property of the graph itself, instead of a sentence buried in a system prompt that the model may or may not obey. Testable transitions turn agent behavior into something a pipeline can check instead of something a human has to eyeball in a chat log.
The Claude Agent SDK's hooks system is the clearest working example of this pattern. It intercepts agent behavior at named points: before a tool call fires (PreToolUse), after it completes (PostToolUse), when the agent finishes responding (Stop), and when a turn errors out (StopFailure). Those are the hooks-as-state-machine-boundary pattern, built directly into a provider SDK.
The five lifecycle stages
Every production agent moves through the same five stages: initialization, input, memory, decision, and execution. Each one is an independent failure surface, and a state machine has to handle each of them explicitly rather than trusting the model to muddle through.
Initialization comes first. This is where the agent gets instantiated: the system prompt loads, tools register, memory context gets seeded, session identity gets established. Skipping a step here produces a subtle failure. If tool registration is incomplete or a schema is malformed, the model can hallucinate a call to a tool that was never actually wired up, and nothing catches it because there's no error surface watching for that specific failure. The state-machine fix is to treat initialization as its own named state, with entry logic and a guard that validates the environment before anything transitions to "ready."
Input comes next: a message, a scheduled trigger, a webhook, a handoff from another agent. The instructive failure mode here is identity. It's identity. An agent whose session state is tracked per channel rather than per person will treat the same user as three different strangers if that user moves from Telegram to Slack to a web client. Fixing that means keying session state to a user_id that's normalized across channels, not to whichever channel_id happened to deliver the message.
Memory is where the agent retrieves whatever context it needs: prior decisions, preferences, task state, domain knowledge. Mem0's State of AI Agent Memory 2026 report found that context drift or retrieval failure during multi-step reasoning was the leading cause of enterprise AI failures in 2025. Running out of room in a context window turned out to be a smaller problem than retrieving the wrong thing while there was still plenty of room left. A three-layer memory system addresses this: vector retrieval for fuzzy, semantic knowledge, key-value stores for structured state like task progress or configuration, and episodic logs for audit trails. Structured state belongs in a key-value store specifically because the agent already knows the exact key it needs. Reaching for a vector search when a direct lookup would do is how drift creeps in.
Decision is where the model reasons over its current state and available tools and proposes what happens next. Without a guard on this transition, a model that picks a tool outside its authorized scope, or proposes a move the graph was never built to permit, just... executes. Silently. The Amazon Kiro incident is the sobering example here: an agentic coding tool deleted and recreated a customer-facing environment because human-approval controls weren't enforced at the transition level. A guard on that specific transition would have blocked it pending sign-off. Business rules belong on the transition, inside the state machine, instead of as a polite request typed into the model's prompt.
Execution is where side effects actually land: files get written, APIs get called, messages go out to real people. Partial execution followed by a crash leaves the world in an inconsistent state with no record of what finished and what didn't. The fix is checkpoint persistence at each state boundary: define the state machine, save a checkpoint at every boundary, let the agent sleep through idle stretches (waiting on a human approval, an external API, whatever), then wake it up exactly where it left off. That pattern applies directly to invoice disputes, procurement approvals, and compliance audits: workflows that don't resolve in one sitting but stretch across hours or days.
Lifecycle hooks as the interface between the state machine and the LLM
Hooks are where the state machine takes control back from the model, and how carefully they're designed determines how bounded an agent's behavior stays once it's live.
A hook is a named point inside the agent's execution where the harness can step in: before a tool call fires, after a response comes back, when an error gets raised. At that point, the harness can inspect what's about to happen, modify it, log it, or stop it.
Without hooks, whatever the model outputs is final. The state machine can watch what happened after the fact, but it has no way to step between the model's decision and the real-world consequence of that decision. With hooks in place, the harness gets to enforce policy before the side effect lands: check authorization, validate that the tool call matches its schema, write the action to an audit log, or demand human approval before anything with real blast radius goes through.
The Claude Agent SDK builds this in directly, with interception points at pre-tool-call (PreToolUse), post-tool-use (PostToolUse), and failure (PostToolUseFailure/StopFailure). The OpenAI Agents SDK takes a related primitive with its handoff abstraction: it names the transition between agents explicitly, which is a form of lifecycle boundary.
Three hooks earn their place in essentially any production agent because they enforce the same lifecycle boundaries the framework primitives expose, regardless of which framework sits underneath.
- A pre-action hook intercepts every tool call before it runs: it validates the schema, checks that the call falls inside the agent's authorized scope, and logs what's about to happen. This is exactly the gate that was missing in the Kiro incident, the "get a second opinion before doing something destructive" checkpoint.
- A state-persistence hook fires after every transition and writes a checkpoint to durable storage, so a crash or an interruption doesn't erase progress; the agent picks up from that exact point instead of restarting from zero.
- An error-boundary hook catches exceptions at the execution layer and routes them into a named error state, rather than letting them bubble up as an unhandled crash. The agent lands in a degraded but recoverable state instead of dying mid-task.
How the major 2026 frameworks implement these primitives
Every major framework shipping in 2026 has converged on the same primitive set: states, transitions, hooks, persistence, error handling. Where they differ is how explicitly each one exposes those primitives, and that difference determines what an engineering team can actually debug and govern once the agent is in production.
LangGraph, available in Python and TypeScript under an MIT license, is the most explicit implementation of the state-machine idea itself: nodes are states, edges are transitions, and the whole graph can be inspected directly. It's become something close to the production standard for stateful, auditable, durable workflows, with deployments running at Klarna, Uber, LinkedIn, BlackRock, JPMorgan, and Replit. OpenSRE, an open-source agent built for incident investigation, runs on top of it.
Microsoft's Agent Framework 1.0, covering Python and.NET under MIT and reaching general availability on April 2, 2026, treats human-in-the-loop as a first-class primitive rather than an afterthought. The approval gate lives in the framework itself, not in a prompt someone hopes the model reads carefully, with responsible AI guardrails available through Azure AI Foundry. For teams where the approval-transition pattern is a compliance requirement rather than a nice-to-have, that's the framework built around exactly that need.
The Claude Agent SDK, in Python and TypeScript, has the most explicit lifecycle-interception API of any provider-native SDK: pre-tool, post-response, and on-error hooks, paired with the deepest MCP integration on the market. Session management uses system-generated session IDs, resumable by the ID that was captured, though the SDK client doesn't support supplying a custom ID. It's Claude-first by design, though it can be redirected to other providers through a compatible gateway, a detail that matters for any team that wants to keep its model options open.
The OpenAI Agents SDK, Python and TypeScript under MIT, centers on the handoff abstraction that names transitions between agents explicitly. Its Skills and Compaction primitives, introduced in an April 2026 update, handle versioned behavior and context management as platform-level sandbox services rather than core SDK primitives. The SDK is tightly scoped, which makes it a strong fit for clean delegation chains and a less natural fit for complex, cyclic graphs.
Google's ADK spans Python, TypeScript, Java, Go, and Kotlin, under Apache 2.0, with 1.0 releases for Java and Go landing in early 2026. It's built around hierarchical multi-agent structures with debugging UIs built in, and that language breadth is a differentiator for enterprise teams working across a polyglot stack that's GCP-native.
It's likely the fastest route to a working prototype of the frameworks listed here, though it's less explicit than LangGraph about how state transitions themselves are represented.
Mastra, a TypeScript framework under an Apache 2.0 core with a source-available enterprise tier, bundles workflows, memory, and a Studio environment into one package. For teams that want lifecycle primitives available out of the box rather than assembled from three or four separate libraries, that bundling is the appeal.
Beneath all of these frameworks runs an agent operating system that provides the runtime environment under the framework itself, handling durable state, scheduling, retries, and execution isolation at the infrastructure level. That's the argument for a purpose-built agent OS instead of cloud tooling retrofitted for the job. It's also the subject of active research, including a SOSP 2026 workshop focused specifically on primitives, isolation models, and scheduling for exactly this layer.
None of these frameworks operate in isolation from each other, either. MCP, donated to the Linux Foundation's Agentic AI Foundation in December 2025 and now backing more than 200 server implementations, and A2A, which absorbed ACP under the same foundation, are the protocols that let agents built on different frameworks share tools and hand work off to each other. A tool built once as an MCP server becomes available to any framework that speaks MCP, which decouples the work of building a tool from the choice of which harness runs it.
Persistence and durable state: the primitive that turns an agent into a long-running system
Without durable state, every agent is stateless at the infrastructure layer, no matter how sophisticated its reasoning looks in a demo. It can simulate memory for as long as a context window holds together, but a context window doesn't survive a crash, a restart, or a process that needs to pause for hours while waiting on a human to approve something.
That's the actual gap between a demo agent and a production one. A demo agent runs start to finish in one sitting, so nothing ever needs to persist. A production agent, handling an invoice dispute that needs several approvals across channels, has to survive the gaps in between without treating the same user as different entities in each channel. It has to write a checkpoint at each state boundary, go quiet during the idle stretches, and wake up exactly where it left off when the next input arrives.
Build the state machine first. Persistence makes that state machine durable enough to run for as long as the actual work takes, beyond the length of one conversation.
Sources
- 6 categories AI Agent Behavior-Improvement Lifecycle | by Chier Hu | Medium
- AI Agent Frameworks (2026 Update): 8 SDKs Compared + the Claude Agent SDK Primitive Reference
- The Architecture of Agency: A Deep Technical Guide to Agentic AI Systems in 2026 | by NJ Raman | Medium
- AI Agent State Machines: Why Explicit Workflows Beat Prompt Loops
- AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
- AgenticOS 2026: Operating Systems Design for AI Agents
- Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents


