Regression Detection in Iterative Agent Development
Four moving targets break agents; catching regressions requires rethinking how tests work.

It doesn't come from one broken line of code. It comes from four things moving at once, prompts, tools, model versions, and orchestration logic, any of which can shift on its own or in combination with the others. Regression testing was built to solve a different problem, and that difference matters for understanding what actually catches it.
Why agents break in four directions at once
Traditional regression testing rests on a quiet assumption: the spec holds still. One variable changes, a human or a test suite checks it, the process moves in a straight line from change to verification to confidence. Agent systems don't offer that kind of stability. A prompt gets rewritten. A tool gets swapped. A model gets upgraded. An orchestration layer gets restructured. Any of these can happen alone, or two or three can happen in the same week, and the interactions between them don't add up neatly. A model upgrade can change how output gets formatted, and that can quietly break a downstream tool-call parser, even if nobody touched the tool itself.
AgentAssay, a framework described by Varun Pratap Bhardwaj in March 2026, starts from a blunt observation: no principled method yet exists for confirming that an agent hasn't regressed after a change to its prompts, tools, models, or orchestration logic. No gap in tooling maturity will close that on its own with better dashboards; it's a structural feature of how these systems are built. No matter what else about the agent changed, every model upgrade is a live regression risk.
The SDAD framework, laid out by Nguyen and Nguyen in May 2026, adds another layer to this. Coding agents now reason through many steps before they produce an answer, working across contexts that run from hundreds of thousands of tokens into the millions. At that scale, the quality of the specification an agent works from directly shapes how faithfully it executes. A spec that's vague, or one that's drifted from what the system actually does, doesn't just produce a slightly worse answer. It compounds, step by step, into a regression that's nearly impossible to trace back to its source.
A fifth dimension sits underneath the four, produced by multi-agent pipelines running several agents at the same time. When several agents run at the same time, their outputs can each look correct in isolation while the interaction between them breaks something downstream. No single agent did anything wrong. The regression lives in the handoff, not in either agent's code.
Keeping pace with agent-speed development
As code merges at agent speed, the regression suite becomes the primary check on what ships, no longer just a backstop behind human review. Hand-maintained test suites were never built to carry that weight alone.
Autonoma's analysis lays out four traits that separate agent-generated code from human-generated code, and each one matters to how regressions get caught or missed. Agent changes have a wider blast radius: a single pull request from an agent tends to touch more of the codebase, per unit of actual intent, than the equivalent change made by a person. At the volume agents produce, human review collapses into pattern-matching rather than real scrutiny, so the test suite ends up doing the reviewing for a large share of what gets merged. Agents also refactor constantly, renaming attributes, restructuring components, and moving elements around, which makes selector-based tests break as a matter of routine. And coverage debt piles up faster than anyone can write tests for it: a model where QA writes tests after features ship guarantees an infinite backlog the moment agents stop pausing for coverage to catch up.
Autonoma also names the cost of getting this wrong: the "merge tax." Every hour a valid pull request sits blocked by a false regression failure has a measurable cost, and a CI gate has to be built to minimize that tax, not just to maximize how much of the code it covers. A suite that throws false positives at scale doesn't just waste time. It teaches engineers to stop trusting failures at all, so real failures start getting ignored along with the noise.
Think of something like Playwright running hand-maintained selectors, and you can picture the failure mode. A selector breaks, CI turns red, and the pull request sits blocked, but the agents generating code don't pause and wait for a fix. They branch again. By the time someone patches the broken selector, several more pull requests have touched the same component tree, and the problem has compounded.
The scale this reaches in practice is documented in a case study from Davis and colleagues at Purdue, "Cheap Code, Costly Judgment." Over twelve weeks, one engineer steering frontier AI coding agents produced 420 thousand lines of production code and over a million lines of tests and supporting tooling. The governance challenge in that project wasn't generating the code. It was keeping the feedback loops inspectable and correctable at that volume.
One might argue the fix is just running more tests, faster. But running more of the wrong kind of test at higher volume doesn't produce more signal. It produces more false positives and more noise. The limitation is architectural.
Binary pass/fail is the wrong verdict for a stochastic system
Even a perfectly maintained test suite runs into a deeper problem: agents don't produce a single output. The same prompt, the same tools, the same model version produces a distribution of possible outputs, not one fixed answer. A binary pass/fail check was built for systems where sameness in equals sameness out. Agents don't work that way, so pass/fail ends up answering a question the system was never built to ask.
That mismatch cuts in both directions. A single run that fails an assertion might sit entirely within the normal variance of an agent that hasn't regressed at all, a false alarm. And a real regression, one that genuinely shifts the distribution of outputs toward worse behavior, can slip past a single-run check if the new distribution still occasionally lands on an output that happens to pass.
AgentAssay replaces the binary verdict with something closer to how a scientist would treat uncertain evidence: three-valued verdicts, PASS, FAIL, or INCONCLUSIVE, grounded in statistical hypothesis testing. An INCONCLUSIVE result is an honest answer when the evidence collected so far isn't enough to say anything with confidence, which is exactly the right thing to tell an engineer in that situation.
To decide how much evidence is enough, AgentAssay uses Sequential Probability Ratio Testing, or SPRT. In practice, SPRT means the test keeps running trials until the evidence crosses a confidence threshold in either direction, then stops, rather than running some fixed number of trials regardless of how much variance shows up along the way. So the posture is meaningfully different than running ten trials just because ten was the number someone picked last year.
The payoff shows in AgentAssay's own experiments across thousands of trials: behavioral fingerprinting, discussed in the next section, reaches 86% detection power in exactly the scenarios where binary pass/fail testing has close to none. Binary testing was built to answer a different kind of question than the one a stochastic agent requires.
SDAD makes a parallel move at the specification layer. Nguyen and Nguyen introduce a metric called Spec Fidelity, because story-point-style binary completion checks can't tell you whether an agent's output actually reflects what the specification intended. Fidelity, in their framing, is a distribution to be measured, not a box to be checked.
Dependency-aware impact analysis catches regressions before the commit lands
If pass/fail verdicts come too late and answer the wrong question, the next move is to push detection earlier, before the commit even lands. An agent that can reason about which tests its own proposed change is likely to affect can run those tests, see the result, and self-correct inside the same execution context, rather than waiting for CI to surface a failure after the fact.
A dependency map connecting source code to tests is the mechanism behind this. When an agent proposes a patch, it checks that map, identifies the subset of tests relevant to the change, runs them, and adjusts before committing anything. Described in the research context as Test-Driven Agentic Development, or TDAD, this pre-change impact analysis is built specifically for AI coding agents, using a dependency graph that scopes verification to the part of the test surface that actually matters and so avoids re-running the whole suite for every change.
Autonoma's coverage approach works on the same principle but from a different angle. Coverage gets re-derived from the structure of the code itself rather than from selectors a human wrote and has to keep updating. It stays current as agents refactor without needing a person to patch it after every commit.
The wider blast radius discussed earlier makes this more than a convenience. Because agent pull requests touch more of the codebase per unit of intent than human ones do, the dependency graph has to understand code structure, not just track imports. If a dependency tracker only follows import statements, it misses the cross-module interactions that agent-driven refactors produce routinely.
The Davis et al. governance model from Purdue offers a useful way to think about why this layer exists at all. Agentic implementation velocity surfaces recurring structural failure classes, and durable governance mechanisms get discovered by studying those failures directly. Pre-commit impact analysis is one such mechanism, and it exists because post-commit CI alone wasn't enough to govern the pace at which agents work.
None of this works, though, if the specification an agent is coding against is vague. SDAD formalizes this cost as the "Ambiguity Tax," the exponential price of working from an unclear spec in agentic development. Dependency analysis operates on intent, and that intent has to be machine-readable to be traced reliably. An ambiguous spec doesn't just produce a worse outcome. It produces dependencies that can't be mapped.
Behavioral fingerprinting as the detection method for regressions that pass/fail cannot see
An agent's behavior is more than the final answer it hands back. It's the sequence of tool calls it makes along the way, the branches it takes and doesn't take, the memory it accesses, the intermediate states it passes through before arriving at an output. Two agents can land on identical final answers while getting there through completely different paths, and one of those paths might be fragile in a way the final answer never shows you.
That gap, between what an agent outputs and how it got there, is what behavioral fingerprinting is built to close. AgentAssay implements it as one of eight technical contributions in Bhardwaj's framework: execution traces get mapped to compact vectors, and those vectors get compared statistically, across many dimensions at once, between one version of an agent and the next. When the comparison is run this way, regressions that are completely invisible at the output level become visible in how the behavior shifted.
The strongest evidence for why this matters is the detection gap itself. AgentAssay's 2026 experiments found behavioral fingerprinting hits 86% detection power in precisely the scenarios where binary pass/fail testing detects almost nothing. Those are regressions that the previous generation of tooling simply couldn't see, because output-level checks were never positioned to see them.
A structurally similar idea appears in the TheBotCompany case, offered here only as background: a verification layer, built by Apollo, catching regressions at the level of a planning cycle, sitting above the unit-test gate. It works because it reasons from execution-level evidence rather than from output-level assertions, which puts it in a position to catch what unit tests were always going to miss.
Fingerprinting also extends naturally into production. AgentAssay's trace-first approach runs its analysis on execution traces that are already being logged, so regression detection on live production runs adds no marginal inference cost. There's no new instrumentation to build. The data was already there; the technique just reads it differently. If you run agents through tools like Codex CLI, you can wire this into existing session hooks, such as the Stop or SessionEnd events, and get automated per-run regression detection without touching the agent's actual execution path.
Memory deserves attention on its own as a sub-case. A regression in how an agent stores, retrieves, or carries context across sessions can leave any single output completely unchanged even as the execution trace underneath shows different retrieval calls, different context assembly, and different session state. Output-level testing cannot see this. The MOMENTO benchmark exists to surface exactly this blind spot, evaluating agents across persistent, multi-session environments where context carried from one session to the next is treated as a first-class input. Existing single-session benchmarks have no mechanism for catching a regression that only appears across sessions.
A purpose-built regression pipeline end to end
None of these three techniques, dependency-aware scoping, statistical verdicts, behavioral fingerprinting, works as a standalone fix. Each one catches a different class of regression, and together they form a pipeline where each stage narrows what the next stage has to deal with.
The first layer sits before the commit. The pipeline uses a code-structure-aware dependency graph to find which tests fall within the blast radius of a proposed change, runs only those tests, and feeds the results back into the agent's own context so it can correct course immediately. This is where the merge tax gets controlled at the source: a change with narrow impact doesn't need to trigger a full suite run, so valid pull requests don't sit blocked waiting on tests that have nothing to do with what actually changed.
The second layer sits at merge. Instead of a binary pass/fail gate, this layer runs stochastic three-valued verdicts backed by SPRT, blocking only on a genuine FAIL and routing INCONCLUSIVE results to a human for a judgment call, rather than halting the pipeline over a single run that might just be noise. Schema validation acts as a fast first filter here: a prompt change that alters output structure fails a schema test immediately, catching an entire class of regression cheaply before the heavier statistical layer even needs to run. In practice, this can be as direct as a pass-rate threshold inside an eval runner, where a non-zero exit code fails the CI job and blocks the merge, a pattern that works with tooling available today.
The third layer runs underneath both, on execution traces rather than outputs. Behavioral fingerprinting compares the shape of how an agent behaves, version over version, catching the regressions in tool-call sequencing, memory access, and intermediate reasoning that no assertion on the final answer would ever surface.
Put together, the three layers answer three different questions: what might this change break, is this output distribution actually different from before, and did the agent's behavior shift in ways the output never shows. An agent is a system that generates a distribution of outcomes, not a fixed function, and a regression pipeline built on that premise, rather than on pass/fail, is the only kind built to track it honestly.
Sources
- SDAD: Spec-Driven Agentic Development for the AI-Native SDLC
- [2603.02601] AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
- Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
- TDAD: Test-Driven Agentic Development
- TheBotCompany: Self-Organizing Multi-agent Systems for Continuous Software Development


