DataAIHub Research · Type-B · v1.0

When AI Agents Fail

Finding the first divergence, containing failure, and evaluating recovery

Published
26 September 2026
Last updated
26 September 2026
Research version
v1.0

AI agents are usually evaluated at the end of a run.

  • Did the task succeed?
  • Was the answer correct?
  • Did the expected state get produced?

Those questions remain important. But for a long-running agent, they don't tell the whole story.

An agent can make an incorrect decision early, continue operating on that incorrect state, and eventually produce an acceptable-looking result. It can also encounter a failure, recover successfully, but take unnecessary or unsafe actions before doing so.

Recent research is making this problem increasingly visible. Locating Hidden Failures Makes Long-Horizon Agents More Reliable (Traverse; arXiv:2609.17930) finds that after a first mistake an agent often fails to recover and rarely catches the error itself, so the run continues, and that detecting whether a trajectory failed is substantially easier than locating the first mistake.

The important question therefore becomes:

When an agent run goes wrong, where did it first become unreliable, what happened afterward, and how well did the system recover?

This article focuses on that part of agent reliability.

The final answer hides the failure

Consider a simplified trajectory:

Simplified agent trajectory
  1. Task
  2. Plan
  3. Search
  4. Interpret result
  5. Choose tool
  6. Execute
  7. Final answer

Now suppose the interpretation is wrong:

Wrong interpretation
  1. Interpretation ✗
  2. Wrong decision
  3. Wrong tool
  4. Wrong action
  5. Final answer

The final answer tells us that something went wrong. It does not necessarily tell us where the trajectory first became wrong.

This distinction matters because later errors may simply be consequences of an earlier one.

DataAIHub's Agent Evaluation guide already covers outcome and trajectory evaluation. The question here goes one level deeper:

Can we locate the point at which the trajectory first diverged?

Recent work suggests that this is difficult even for strong automated judges. Traverse studies 2,518 agent trajectories and classifies 6,967 mistakes. On the long software-engineering and computer-use runs, the strongest of six frontier judges locates the first mistake in fewer than one-third of those runs. That localization result does not apply uniformly to all 2,518 trajectories: exact localization is higher on the short science runs in the same study.

The same study records a limit on reading a run from its ending. 343 of the 2,518 runs (14%) were solved despite a flagged mistake, so a successful final outcome can still hide a problematic intermediate trajectory.

That makes failure localization a problem in its own right.

Find the first divergence

A long trajectory can contain many incorrect steps:

First divergence in a trajectory
  1. Understand task — correct
  2. Search — correct
  3. Inspect result — correct
  4. Misinterpret result — first divergence
  5. Select wrong tool — consequence
  6. Wrong arguments — consequence
  7. Change system state — consequence
  8. Compensate — consequence
  9. Final answer — failed

A conventional failure report might identify steps 5–9.

A more useful investigation asks whether those steps were caused by step 4.

First divergence

For this article, first divergence means:

The earliest step at which the observed trajectory materially departs from an acceptable trajectory.

This does not mean that there is always one objectively correct sequence of actions. Agents may have many valid ways to complete the same task. Traverse also finds that not every failure has one identifiable first mistake: 588 of the 2,518 runs (23%) failed with no single decisive mistake, a diffuse failure rather than one divergence point.

The goal is instead to identify the earliest point where the trajectory becomes inconsistent with the available evidence, task requirements, system constraints, or a state from which safe recovery remains possible.

That gives us three different things to distinguish:

Divergence, consequences, and final failure
  1. First divergence
  2. Downstream consequences
  3. Final failure

They may all occur at different points.

Why this matters

If step 7 is wrong because of step 4, fixing step 7 alone may only remove one symptom.

A useful investigation therefore asks:

  • What did the agent know at the point of divergence?
  • What did it infer?
  • What action did that inference produce?
  • Which later decisions depended on it?
  • Could the trajectory still have been recovered at that point?

The result is a more useful failure description than simply:

Agent failed at step 9.

Trace the propagation

After a trajectory diverges, the error can spread.

A simplified chain looks like this:

Failure propagation
  1. Bad interpretation
  2. Wrong decision
  3. Wrong action
  4. Changed state
  5. New observation
  6. Next decision
  7. Further action

This creates an important distinction:

Not every later error is a new root cause.

Some are propagated consequences of the original divergence.

Traverse labels each mistake as either a root fault or a cascade, a downstream consequence of an earlier mistake, which is why a long trace is hard to diagnose from a list of later errors alone.

The engineering question is therefore not only:

What failed?

It is:

What did the failure cause afterward?

Consider an agent that misinterprets a retrieved record.

If it only uses that interpretation to draft a response, the failure may remain contained.

If it uses it to select another tool, modify state, and make subsequent decisions based on the modified state, the same initial error has a much larger propagation path.

We can represent that difference as:

Contained versus propagated divergence
Divergence
  • Contained
  • Propagated
    • Decision
    • Tool
    • State

This failure propagation model is a synthesis rather than an established universal framework. Its value is as an engineering lens for reading agent traces.

Instead of treating the trace as a list of independent events, we can ask which events are causally connected.

Evaluate the recovery

A failure does not necessarily mean the run is lost.

An agent may recognize the problem and recover.

But recovery itself can fail.

Compare:

Recovery comparison

Repeated retry

  1. Tool failure
  2. Retry
  3. Same failure
  4. Retry
  5. Same failure

with:

Recovery that changes strategy

  1. Tool failure
  2. Diagnose
  3. Change strategy
  4. Alternative action
  5. Verify
  6. Continue

Both are responses to failure. Only one demonstrates meaningful recovery.

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling (Plan-RewardBench; arXiv:2604.08178) treats robust error recovery as its own trajectory-level category, alongside tool unavailability and planning. Its recovery negatives include blind retries: repeating the same failing call without a meaningful change. The trajectories it prefers diagnose the error and change strategy.

Tool-Reflection-Bench, introduced in Failure Makes the Agent Stronger (ACL Findings 2026), scores miniature trajectories of an erroneous tool call, then a reflection, then a corrected call. The correction itself is what is assessed.

On that evidence, a retry is not yet a recovery. Plan-RewardBench treats a repeated failing call as a negative, and a diagnosis that changes strategy as the preferred recovery. That is a benchmark distinction, not a universal rule.

A recovery investigation can ask:

Recovery investigation
  1. Failure detected
  2. Was the failure recognized?
  3. Was the cause understood?
  4. Did the strategy change?
  5. Was the original error avoided?
  6. Was additional damage prevented?
  7. Was the resulting state verified?

DataAIHub already covers mechanisms such as agent planning, durable execution, and human-in-the-loop.

The question here is different:

Was the recovery behavior itself appropriate?

Contain, recover, escalate, or stop

Recovery is not always the right response.

Once a failure is detected, a production system has several possible directions. For this article, we treat the choice among them as one engineering decision, not as a standard control policy:

Recover, escalate, or terminate
Failure
  • Recover
  • Escalate
  • Terminate

The appropriate choice depends on the failure and the state of the system.

Factors that can matter include:

  • whether the failure is understood;
  • whether the current state is trustworthy;
  • whether the action is reversible;
  • whether external side effects have occurred;
  • whether the same failure is repeating;
  • whether continuing could cause additional damage;
  • whether a human decision is required.

There is no evidence for a universal threshold that determines when an agent should recover, escalate, or terminate. We should therefore treat this as a decision problem, not a fixed algorithm.

Containment

Before asking how to recover, there is another question:

How far can the failure still spread?

Consider:

Contained path
  1. Search
  2. Wrong interpretation
  3. Draft response
  4. Correction

versus:

Propagated path
  1. Search
  2. Wrong interpretation
  3. Wrong decision
  4. Database write
  5. External notification
  6. Changed system state

The second trajectory has allowed the same initial mistake to affect substantially more state.

For this article, failure containment means:

Limit the additional trajectory or external state that a detected failure can affect.

This does not replace guardrails, approval mechanisms, idempotency, or durable execution. DataAIHub already covers those mechanisms in its Guardrails, Tool Calling, Human-in-the-Loop, and Durable Execution guides.

The new question is temporal:

Where in the trajectory should the failure stop propagating?

Turn failures into regression knowledge

A production failure is most valuable when the system learns from it.

The familiar feedback loop is:

Familiar feedback loop
  1. Production run
  2. Failure
  3. Evaluation case
  4. Regression suite
  5. Future change

OpenAI's Evaluate agent workflows documentation describes this kind of trace-driven workflow for its own agents: traces containing model calls, tool calls, guardrails, and handoffs, then datasets and repeatable evaluation runs. That is one documented product workflow, not a claim about every agent-evaluation system.

DataAIHub already covers the broader production trace → evaluation → regression loop.

The question here is what information should survive when a failure becomes a regression case.

Failure-aware case

A simple test might contain:

  1. Input
  2. Expected output

A failure-aware case could preserve more of the run:

  1. Scenario
  2. Observed trajectory
  3. First divergence
  4. Propagation
  5. Recovery behavior
  6. Final state

That makes it possible to test more than the final result:

  • divergence recognized
  • propagation contained
  • recovery appropriate
  • repeated failure avoided
  • final state acceptable

This is a proposed engineering model, not an established industry standard.

Its purpose is practical: preserve enough information about a failure that a future model, prompt, tool, or orchestration change can be tested against the behavior that actually caused the incident.

For example:

Failure-aware regression

Traditional regression

  1. Input
  2. Expected output
  3. PASS

Failure-aware regression

  1. Input
  2. Failure at step 4
  3. Expected containment
  4. Expected recovery
  5. Acceptable final state

A system can therefore fail the regression even when it eventually produces the right final answer—if it repeats an unacceptable intermediate behavior.

Reliability becomes a feedback process

Putting the pieces together, this article reads agent reliability as its own engineering lens: a feedback process rather than an end-of-run score.

Reliability feedback loop
  1. Agent run
  2. Trajectory
    • Success
    • Divergence
  3. Propagation
  4. Detection
    • Recover
    • Escalate
    • Stop
  5. Failure case
    1. Evaluation
    2. Regression
    3. Change
  6. New run

New run → Agent run

The important change is that evaluation is no longer only an end-of-run activity.

A useful reliability process asks:

  • Where did the trajectory first diverge?
  • How far did the error propagate?
  • Was the failure detected?
  • Was the recovery appropriate?
  • Was the failure contained?
  • What should happen differently on the next run?

As that lens, reliability is a feedback process rather than a property established by a single benchmark score.

Key takeaways

  1. Final outcome is necessary but insufficient. A successful final state does not necessarily mean the trajectory was reliable.
  2. Find the first divergence. Later errors may be symptoms of an earlier mistake.
  3. Trace propagation, not just failure. Understanding which later decisions depended on the original divergence makes debugging more useful.
  4. Evaluate recovery itself. Retrying is not necessarily recovering. A useful recovery changes behavior, avoids repeating the cause, and restores an acceptable state.
  5. Containment matters. The consequence of a failure depends partly on how far it is allowed to propagate.
  6. Preserve failure structure in regression tests. A useful failure case may need to remember not just the final state, but the divergence, propagation, and expected recovery.

Research perspective

Traverse, Plan-RewardBench, and Tool-Reflection-Bench do not yet provide a single accepted framework for all of these concepts. First-error localization, trajectory evaluation, recovery evaluation, and long-horizon reliability are active areas of research.

The useful synthesis for engineering is narrower:

A production agent should not be evaluated only on whether it eventually succeeds. When a trajectory fails, engineers need to know where it first diverged, how the failure propagated, whether the system recovered appropriately, and what should be captured for the next evaluation cycle.

That gives us a practical chain:

Divergence → Propagation → Recovery → Learning

And that chain is where the next generation of agent reliability work is likely to become more operationally useful.