DataAIHub Research · Type-B · v1.0
When AI Agents Fail
Finding the first divergence, containing failure, and evaluating recovery
- Published
- 26 September 2026
- Last updated
- 26 September 2026
- Research version
- v1.0
AI agents are usually evaluated at the end of a run.
- Did the task succeed?
- Was the answer correct?
- Did the expected state get produced?
Those questions remain important. But for a long-running agent, they don't tell the whole story.
An agent can make an incorrect decision early, continue operating on that incorrect state, and eventually produce an acceptable-looking result. It can also encounter a failure, recover successfully, but take unnecessary or unsafe actions before doing so.
Recent research is making this problem increasingly visible. Locating Hidden Failures Makes Long-Horizon Agents More Reliable (Traverse; arXiv:2609.17930) finds that after a first mistake an agent often fails to recover and rarely catches the error itself, so the run continues, and that detecting whether a trajectory failed is substantially easier than locating the first mistake.
The important question therefore becomes:
When an agent run goes wrong, where did it first become unreliable, what happened afterward, and how well did the system recover?
This article focuses on that part of agent reliability.
The final answer hides the failure
Consider a simplified trajectory:
- Task
- Plan
- Search
- Interpret result
- Choose tool
- Execute
- Final answer
Now suppose the interpretation is wrong:
- Interpretation ✗
- Wrong decision
- Wrong tool
- Wrong action
- Final answer
The final answer tells us that something went wrong. It does not necessarily tell us where the trajectory first became wrong.
This distinction matters because later errors may simply be consequences of an earlier one.
DataAIHub's Agent Evaluation guide already covers outcome and trajectory evaluation. The question here goes one level deeper:
Can we locate the point at which the trajectory first diverged?
Recent work suggests that this is difficult even for strong automated judges. Traverse studies 2,518 agent trajectories and classifies 6,967 mistakes. On the long software-engineering and computer-use runs, the strongest of six frontier judges locates the first mistake in fewer than one-third of those runs. That localization result does not apply uniformly to all 2,518 trajectories: exact localization is higher on the short science runs in the same study.
The same study records a limit on reading a run from its ending. 343 of the 2,518 runs (14%) were solved despite a flagged mistake, so a successful final outcome can still hide a problematic intermediate trajectory.
That makes failure localization a problem in its own right.
Find the first divergence
A long trajectory can contain many incorrect steps:
- Understand task — correct
- Search — correct
- Inspect result — correct
- Misinterpret result — first divergence
- Select wrong tool — consequence
- Wrong arguments — consequence
- Change system state — consequence
- Compensate — consequence
- Final answer — failed
A conventional failure report might identify steps 5–9.
A more useful investigation asks whether those steps were caused by step 4.
First divergence
For this article, first divergence means:
The earliest step at which the observed trajectory materially departs from an acceptable trajectory.
This does not mean that there is always one objectively correct sequence of actions. Agents may have many valid ways to complete the same task. Traverse also finds that not every failure has one identifiable first mistake: 588 of the 2,518 runs (23%) failed with no single decisive mistake, a diffuse failure rather than one divergence point.
The goal is instead to identify the earliest point where the trajectory becomes inconsistent with the available evidence, task requirements, system constraints, or a state from which safe recovery remains possible.
That gives us three different things to distinguish:
- First divergence
- Downstream consequences
- Final failure
They may all occur at different points.
Why this matters
If step 7 is wrong because of step 4, fixing step 7 alone may only remove one symptom.
A useful investigation therefore asks:
- What did the agent know at the point of divergence?
- What did it infer?
- What action did that inference produce?
- Which later decisions depended on it?
- Could the trajectory still have been recovered at that point?
The result is a more useful failure description than simply:
Agent failed at step 9.Trace the propagation
After a trajectory diverges, the error can spread.
A simplified chain looks like this:
- Bad interpretation
- Wrong decision
- Wrong action
- Changed state
- New observation
- Next decision
- Further action
This creates an important distinction:
Not every later error is a new root cause.
Some are propagated consequences of the original divergence.
Traverse labels each mistake as either a root fault or a cascade, a downstream consequence of an earlier mistake, which is why a long trace is hard to diagnose from a list of later errors alone.
The engineering question is therefore not only:
What failed?
It is:
What did the failure cause afterward?
Consider an agent that misinterprets a retrieved record.
If it only uses that interpretation to draft a response, the failure may remain contained.
If it uses it to select another tool, modify state, and make subsequent decisions based on the modified state, the same initial error has a much larger propagation path.
We can represent that difference as:
- Contained
- Propagated
- Decision
- Tool
- State
This failure propagation model is a synthesis rather than an established universal framework. Its value is as an engineering lens for reading agent traces.
Instead of treating the trace as a list of independent events, we can ask which events are causally connected.
Evaluate the recovery
A failure does not necessarily mean the run is lost.
An agent may recognize the problem and recover.
But recovery itself can fail.
Compare:
Repeated retry
- Tool failure
- Retry
- Same failure
- Retry
- Same failure
with:
Recovery that changes strategy
- Tool failure
- Diagnose
- Change strategy
- Alternative action
- Verify
- Continue
Both are responses to failure. Only one demonstrates meaningful recovery.
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling (Plan-RewardBench; arXiv:2604.08178) treats robust error recovery as its own trajectory-level category, alongside tool unavailability and planning. Its recovery negatives include blind retries: repeating the same failing call without a meaningful change. The trajectories it prefers diagnose the error and change strategy.
Tool-Reflection-Bench, introduced in Failure Makes the Agent Stronger (ACL Findings 2026), scores miniature trajectories of an erroneous tool call, then a reflection, then a corrected call. The correction itself is what is assessed.
On that evidence, a retry is not yet a recovery. Plan-RewardBench treats a repeated failing call as a negative, and a diagnosis that changes strategy as the preferred recovery. That is a benchmark distinction, not a universal rule.
A recovery investigation can ask:
- Failure detected
- Was the failure recognized?
- Was the cause understood?
- Did the strategy change?
- Was the original error avoided?
- Was additional damage prevented?
- Was the resulting state verified?
DataAIHub already covers mechanisms such as agent planning, durable execution, and human-in-the-loop.
The question here is different:
Was the recovery behavior itself appropriate?
Contain, recover, escalate, or stop
Recovery is not always the right response.
Once a failure is detected, a production system has several possible directions. For this article, we treat the choice among them as one engineering decision, not as a standard control policy:
- Recover
- Escalate
- Terminate
The appropriate choice depends on the failure and the state of the system.
Factors that can matter include:
- whether the failure is understood;
- whether the current state is trustworthy;
- whether the action is reversible;
- whether external side effects have occurred;
- whether the same failure is repeating;
- whether continuing could cause additional damage;
- whether a human decision is required.
There is no evidence for a universal threshold that determines when an agent should recover, escalate, or terminate. We should therefore treat this as a decision problem, not a fixed algorithm.
Containment
Before asking how to recover, there is another question:
How far can the failure still spread?
Consider:
- Search
- Wrong interpretation
- Draft response
- Correction
versus:
- Search
- Wrong interpretation
- Wrong decision
- Database write
- External notification
- Changed system state
The second trajectory has allowed the same initial mistake to affect substantially more state.
For this article, failure containment means:
Limit the additional trajectory or external state that a detected failure can affect.
This does not replace guardrails, approval mechanisms, idempotency, or durable execution. DataAIHub already covers those mechanisms in its Guardrails, Tool Calling, Human-in-the-Loop, and Durable Execution guides.
The new question is temporal:
Where in the trajectory should the failure stop propagating?
Turn failures into regression knowledge
A production failure is most valuable when the system learns from it.
The familiar feedback loop is:
- Production run
- Failure
- Evaluation case
- Regression suite
- Future change
OpenAI's Evaluate agent workflows documentation describes this kind of trace-driven workflow for its own agents: traces containing model calls, tool calls, guardrails, and handoffs, then datasets and repeatable evaluation runs. That is one documented product workflow, not a claim about every agent-evaluation system.
DataAIHub already covers the broader production trace → evaluation → regression loop.
The question here is what information should survive when a failure becomes a regression case.
A simple test might contain:
- Input
- Expected output
A failure-aware case could preserve more of the run:
- Scenario
- Observed trajectory
- First divergence
- Propagation
- Recovery behavior
- Final state
That makes it possible to test more than the final result:
- divergence recognized
- propagation contained
- recovery appropriate
- repeated failure avoided
- final state acceptable
This is a proposed engineering model, not an established industry standard.
Its purpose is practical: preserve enough information about a failure that a future model, prompt, tool, or orchestration change can be tested against the behavior that actually caused the incident.
For example:
Traditional regression
- Input
- Expected output
- PASS
Failure-aware regression
- Input
- Failure at step 4
- Expected containment
- Expected recovery
- Acceptable final state
A system can therefore fail the regression even when it eventually produces the right final answer—if it repeats an unacceptable intermediate behavior.
Reliability becomes a feedback process
Putting the pieces together, this article reads agent reliability as its own engineering lens: a feedback process rather than an end-of-run score.
- Agent run
- Trajectory
- Success
- Divergence
- Propagation
- Detection
- Recover
- Escalate
- Stop
- Failure case
- Evaluation
- Regression
- Change
- New run
New run → Agent run
The important change is that evaluation is no longer only an end-of-run activity.
A useful reliability process asks:
- Where did the trajectory first diverge?
- How far did the error propagate?
- Was the failure detected?
- Was the recovery appropriate?
- Was the failure contained?
- What should happen differently on the next run?
As that lens, reliability is a feedback process rather than a property established by a single benchmark score.
Key takeaways
- Final outcome is necessary but insufficient. A successful final state does not necessarily mean the trajectory was reliable.
- Find the first divergence. Later errors may be symptoms of an earlier mistake.
- Trace propagation, not just failure. Understanding which later decisions depended on the original divergence makes debugging more useful.
- Evaluate recovery itself. Retrying is not necessarily recovering. A useful recovery changes behavior, avoids repeating the cause, and restores an acceptable state.
- Containment matters. The consequence of a failure depends partly on how far it is allowed to propagate.
- Preserve failure structure in regression tests. A useful failure case may need to remember not just the final state, but the divergence, propagation, and expected recovery.
Research perspective
Traverse, Plan-RewardBench, and Tool-Reflection-Bench do not yet provide a single accepted framework for all of these concepts. First-error localization, trajectory evaluation, recovery evaluation, and long-horizon reliability are active areas of research.
The useful synthesis for engineering is narrower:
A production agent should not be evaluated only on whether it eventually succeeds. When a trajectory fails, engineers need to know where it first diverged, how the failure propagated, whether the system recovered appropriately, and what should be captured for the next evaluation cycle.
That gives us a practical chain:
Divergence → Propagation → Recovery → Learning
And that chain is where the next generation of agent reliability work is likely to become more operationally useful.