The hidden ambiguity in a tool call
AI agents are usually described as a loop: the user makes a request, the agent selects a tool, the tool returns a result, and the agent continues from that result.
- User
- Agent
- Tool
- Result
- Agent
For a read-only operation, that model is usually adequate. It becomes incomplete once a tool changes something outside the agent—creating an invoice, charging a customer, modifying a database, sending an email, provisioning infrastructure. The problem is then not necessarily that the tool failed. The problem is that the agent may no longer know what happened.
A tool call that changes external state can complete even when its response never reaches the agent. In the sequence below, the tool commits the transaction and the invoice exists in the database, but the response is lost and the agent observes only a timeout.
- RequestAgent
create_invoice(...) - ExecutionToolcommit transaction
- ExecutionDatabaseinvoice created
- External state changedExternal effect exists
- ResponseTool responselost / timeout
- Agent observationAgent observes
TIMEOUT
TIMEOUTThe timeout does not establish that the invoice failed. It establishes something narrower: from the agent's position, the outcome is unknown. That distinction is the foundation of a class of problems that becomes increasingly important as AI agents move from retrieving information to changing the external world.
Tool execution is not necessarily an atomic operation. A simple application often treats a function call as something that either succeeds or fails, but an agent tool invocation crosses several potentially failing boundaries: the network, the tool service, the database or external API that applies the side effect, and the path that carries the result back.
call()returnssuccessorfailureBefore the tool receives the request
- Agentrequest
- Network
A failure here is good evidence that this attempt did not cause the effect.
While the request is processed
- Tool service
- Database / API / external system
- Side effect
After the external state has changed
- Responsereturns
- Agent
A failure here leaves the agent unable to infer external state from the timeout.
A failure can occur before the tool receives the request, while the request is being processed, after the external state has already changed, or while the result is travelling back to the agent. These situations are not equivalent. A request that never reaches the tool leaves good evidence that this attempt did not cause the effect; a response lost after the database commits leaves the agent unable to infer the state of the external system from the timeout at all.
This failure mode has recently received explicit attention in research on LLM agents. A 2026 study describes common agent frameworks as implicitly treating tool calls as atomic success/failure operations, while real systems can experience timeout-after-dispatch, delayed visibility, and partial state updates. The authors evaluate postcondition verification, verify-before-retry, and idempotency keys as mechanisms for handling these failures.
TIMEOUT does not mean FAILURE
An agent that executes create_invoice(customer=42, amount=100) and receives TIMEOUT is facing at least two possible realities. In one, the tool failed before creating the invoice. In the other, the invoice was created and only the response was lost. The agent's observation is identical in both; the external state is not.
Reality A
- Request
- Tool
- Failed before creating invoice
External effect: absent
Reality B
- Request
- Tool
- Invoice created
- Response lost
External effect: invoice exists
TIMEOUT≠FAILURE
TIMEOUT → OUTCOME UNKNOWN
Retrying an operation whose outcome is unknown is fundamentally different from retrying a confirmed failure, and most of the rest of this analysis follows from treating the two differently.
When a retry can create a second effect
An agent that receives a timeout may reasonably decide, "The invoice creation failed. I'll try again." If the first attempt had in fact committed, the retry creates a second invoice. The agent believes it created one; the external system now contains two.
Attempt 1
create_invoice()- Invoice created
- Response lost
TIMEOUT
Attempt 2 (retry)
create_invoice()- Second invoice created
SUCCESS
External system
INV-001$100INV-002$100
Agent believes
One invoice was created
The retry was a sensible action for an agent that believed the first operation had failed. What failed was the assumption underneath it—that timeout → no side effect. It wasn't valid.
Idempotency and operation identity
Idempotency is one common way to make such retries safe, and one common mechanism is to give the logical operation an identity such as operation_id = OP-123. The first attempt creates the invoice and records its result against that identity. If the response is lost and the operation is retried, the tool recognizes OP-123 as already processed and returns the original result instead of executing again.
OP-123OP-123- Create invoice
INV-001
OP-123- Already processed
- Return original result (INV-001)
The property this provides is stronger than "don't execute this function twice":
Repeated attempts representing the same logical operation converge on the same durable outcome.
A tool may legitimately receive the same operation more than once. What must not happen twice is the logical external effect: repeated attempts of the same operation should converge on a single durable effect.
Operation identity is not the same as request identity
There is a subtler problem. If OP-123 with amount = 100 succeeds and a later retry arrives as OP-123 with amount = 500, the system should not return the original result. The identifier refers to a logical operation whose intent must remain consistent, so it has to be bound to an intent fingerprint as well as to the result it produced.
Recorded operation
- operation_id
OP-123- intent fingerprint
customer=42, amount=100- result
INV-001
Later request
- operation_id
OP-123- intent fingerprint
customer=42, amount=500
A request that reuses the identifier with a different intent should produce a conflict rather than silently executing a different operation. Robust idempotency therefore requires more than remembering a key; it requires associating the key with the intent and resulting operation.
Verification is useful, but not definitive
Idempotency is not always available. When an external API does not support an idempotency key, the system can respond to a TIMEOUT by establishing the postcondition—whether the operation associated with OP-123 had already created an invoice—before deciding whether to retry. That step is not another execution. It is an observation of external state, and it changes the execution model from retry-on-error to verify-then-decide.
- Execute
- Outcome unknown
- Verify external state
- Already happenedRecover
- Did not happenPotentially retry
- UnknownReconcile
The 2026 study mentioned earlier specifically evaluates this verify-before-retry approach and reports reduced duplicate actions under injected non-atomic failures.
Verification is not magic, however, because the verification system has its own consistency semantics. If the tool writes to a primary database and the verification query reads from a replica, the write can be committed while the check still returns "not found". "Not found" does not necessarily prove that the write never happened, so verification may need to produce three outcomes rather than a Boolean.
Write path
- Write
- Primary
- Committed
Verification path
- Verify
- Replica
- Not found
NOT FOUND≠proof of absence
This is another reason UNKNOWN belongs in the model as a first-class state. A reliable system should not turn uncertainty into an arbitrary decision merely because its API expects a Boolean.
Partial execution creates another problem
Some tools perform multiple state changes. A create_order() tool might create an order, reserve inventory and send a confirmation. If the first two steps succeed and the confirmation fails, the tool returns ERROR—but calling that simply FAILURE hides useful information. The external world is now in a partially changed state, and a retry of the entire operation may create another order or another reservation.
Execution
- create order(completed)
- reserve inventory(completed)
- send confirmation(failed)
Tool returns ERROR
Resulting external state
- order exists
- inventory reserved
- confirmation absent
Recovery
- Reconcile
- Determine completed effects
- Continue / compensate / escalate
The correct recovery strategy could instead be to reconcile: determine which effects completed, then continue, compensate or escalate. This is closely related to long-standing distributed-systems techniques such as idempotency, reconciliation and compensating transactions. What changes with AI agents is that an LLM-driven runtime may otherwise interpret a generic tool error as an instruction to simply try again.
Tool semantics are claims, not guarantees
Not all side effects are equivalent, and several properties of a tool's effects determine what the runtime can safely do after an ambiguous outcome.
| Dimension | Examples | Semantics | Retry implication |
|---|---|---|---|
| Read-only | search_customer()get_invoice()list_projects() | No external state is changed. | Retry is generally much simpler. |
| Idempotent write | set_customer_email(...) | Repeated execution with the same logical intent can converge on the same state. | Repeating the same intent converges rather than accumulating effects. |
| Non-idempotent write | create_invoice(...)charge_customer(...)send_message(...) | Repeated execution can produce additional effects. | A retry after an ambiguous outcome can create a second effect. |
| Destructive | delete_account(...)delete_database(...) | The consequences may be difficult or impossible to reverse. | An incorrect retry may not be reversible. |
| External-world | send_email(...)place_order(...)book_flight(...)publish_message(...) | The side effect may occur outside the system controlling the agent runtime. | The runtime may depend on the external system to establish what happened. |
Tool metadata is useful—but it isn't proof
This problem is becoming particularly relevant to tool protocols such as MCP, which defines tool annotations for declaring a tool's behavior. idempotentHint, for example, indicates that repeatedly calling a tool with the same arguments should have no additional effect on its environment.
| Annotation | Declares |
|---|---|
readOnlyHint | If true, the tool does not modify its environment. |
destructiveHint | If true, the tool may perform destructive updates to its environment; if false, only additive updates. |
idempotentHint | If true, calling the tool repeatedly with the same arguments has no additional effect on its environment. |
openWorldHint | If true, the tool may interact with an "open world" of external entities. |
destructiveHint and idempotentHint are meaningful only when readOnlyHint is false.
But MCP makes an important qualification. The specification states that all properties in ToolAnnotations are hints:
They are not guaranteed to provide a faithful description of tool behavior.
Clients should not treat annotations from untrusted servers as authoritative safety controls. A tool can declare idempotentHint = true without giving us independent proof that repeated execution is actually idempotent.
DECLARED SEMANTICS≠ACTUAL SEMANTICS
This isn't merely a theoretical concern. An issue in the official MCP servers repository reported that the sequentialthinking tool was marked as read-only and idempotent even though it accumulates session state across calls. The reported behavior therefore conflicts with those annotations: repeated calls can change the accumulated state.
Declared (tool annotations)
readOnlyHint = trueidempotentHint = true
Reported behavior (issue #4721)
- Accumulates session state across calls
- Repeated calls can change the accumulated state
The question that matters is not whether a tool has an idempotency annotation, but:
What evidence do we have that the declared semantic property actually holds?
From declarations to evidence
Evidence has different strengths, and there are several ways to learn something about a tool.
| Evidence | What it can tell us |
|---|---|
| Tool schema | Expected inputs and outputs |
| Tool annotation | Declared behavioral properties |
| Documentation | Provider's stated semantics |
| Source code | Some implementation behavior |
| Runtime observation | What happened in a particular execution |
| Postcondition check | Whether a particular state was observed |
| Repeated controlled execution | Behavior under tested conditions |
These should not be treated as equivalent. The same idempotency property can rest on quite different evidence: an MCP annotation of idempotentHint = true is a declaration; a source-code analysis that finds an operation-key lookup is implementation evidence; a controlled experiment in which two equivalent calls produce one durable effect is an observation.
| Source | Finding | Status |
|---|---|---|
| MCP annotation | idempotentHint = true | DECLARED |
| Source-code analysis | Operation-key lookup found | IMPLEMENTATION EVIDENCE |
| Controlled experiment | Two equivalent calls produce one durable effect | OBSERVED |
None of these should automatically be presented as mathematical proof of universal idempotency.
A useful vocabulary for tool semantics
Reducing a tool to SAFE or UNSAFE discards exactly these distinctions. A more useful model reports the status of each semantic property separately, using a vocabulary that records where a claim comes from rather than how much to trust it.
- DECLARED
- Stated in tool metadata or annotations
- DOCUMENTED
- Stated in the provider's documentation
- INFERRED
- Derived from implementation analysis
- OBSERVED
- Seen in runtime or controlled execution
- CONFLICTING
- Available evidence disagrees
- UNKNOWN
- No adequate evidence either way
Applied to create_invoice, that vocabulary produces a profile rather than a verdict. The sequentialthinking case from the previous section would appear in the same form as CONFLICTING: declared read-only and idempotent, with an invocation that changes session state. Either profile is more informative than a generic reliability score.
create_invoiceHypothetical tool, not an analysis of an actual implementation- Writes state
- YESBasis: illustrative assumption
- External effect
- YESBasis: illustrative assumption
- Idempotency
- DECLARED TRUEBasis: hypothetical tool annotation
- Idempotency
- NOT VERIFIED
- Retry semantics
- UNKNOWN
- Verification mechanism
- NOT SPECIFIED
Execution semantics belong in the runtime
The deeper problem is semantic, not specifically about LLMs; the underlying failure mode is familiar from distributed systems. AI agents make the boundary especially important because the system generating the next action may be a probabilistic model. An agent can receive TimeoutError and generate "The operation failed. I'll retry." The runtime should not rely on that interpretation when the external state is uncertain.
- Model
Agent
Proposes action
- Deterministic application logic
Agent runtime
- Operation identity
- Execution policy
- Verification
- Retry semantics
- External system
External tool
The model proposes an action. The application/runtime should own the semantics of executing that action safely.
What a robust tool contract might eventually contain
A function signature tells us what inputs the function accepts. A side-effecting tool could conceptually declare far more: what it changes, how operations are identified, what to do when an outcome is ambiguous, and how its effects can be verified or compensated.
create_invoice(customer_id, amount)Execution contracttool: create_invoice
effects
- writes_state
- true
- external
- true
- destructive
- false
identity
- operation_key
- operation_id
- intent_binding
- required
retry
- ambiguous_outcome
- verify_before_retry
verification
- supported
- true
- method
- get_invoice_by_operation_id
compensation
- supported
- false
The signature does not necessarily tell us what happens if:
- the request times out,
- the side effect commits but the response disappears,
- the same request is retried,
- the retry has different arguments,
- the verification read is stale,
- only part of the operation succeeds.
Those are execution semantics.
What should developer tooling actually know?
If a tool declares idempotent = true, can a developer tool independently find evidence supporting or contradicting that claim? It might combine tool metadata, documentation, source code and runtime observations, and report what each contributes rather than collapsing them into a single answer.
create_invoiceHypothetical scenario, not findings from an actual tool- Declared
- idempotent = true
- Source evidence
- Creates new record using generated invoice ID
- Runtime evidence
- Two equivalent invocations → two records
- Assessment
- CONFLICTING EVIDENCE
The useful output isn't "This tool is unsafe." It is:
"The available evidence conflicts with the declared idempotency property."
Analysis tools should report what they can establish rather than pretending to prove properties they cannot. As agents move from retrieving information toward changing the world, that discipline matters more, and the engineering question shifts from whether the model can call the right tool to:
What guarantees surround the effect produced by that tool call?
A mature agent system therefore needs to reason about at least three separate things.
Intent. What operation did the agent request?
Execution. What happened while attempting it?
External state. What actually changed?
Those three can diverge, and once they do, a generic success/failure result is often insufficient.
Agent intent≠Execution outcome≠Observed external state
Practical design principles
For side-effecting agent tools, several principles follow.
Treat ambiguous outcomes as ambiguous. TIMEOUT ≠ FAILURE. Do not automatically convert uncertainty into failure.
Give logical operations an identity. An operation_id allows retries and reconciliation to refer to the same logical operation.
Bind identity to intent. The same operation key should not silently represent different requests.
Prefer verify-before-retry. When the outcome is unknown, establish external state before creating another effect.
Make verification explicit. A verification mechanism should itself have known consistency properties.
Distinguish partial execution. An operation that performed some effects is not equivalent to an operation that performed none.
Treat tool annotations as claims. Metadata such as MCP's idempotentHint is useful, but it is not independent proof.
Keep execution semantics outside the model. The LLM can propose the action; deterministic application logic should control the consequences of ambiguous execution.
The open tooling question
This leads to an interesting, but still unresolved, question:
Can developer tooling construct an evidence-backed semantic profile of an AI tool without overclaiming what it knows?
A profile like the one shown earlier for create_invoice is a very different product from an "AI reliability score." Its value lies in keeping apart kinds of knowledge that are easily collapsed into one judgement.
- Declaration
- What the developer declared
- Implementation
- What the code appears to do
- Documentation
- What documentation says
- Runtime evidence
- What runtime evidence demonstrates
- Unknown
- What remains unknown
This is potentially useful territory for developer tooling—but it needs to be demonstrated rather than assumed.
The central problem with side-effecting AI tools is not that agents sometimes make mistakes. It is more fundamental:
The agent can lose knowledge about the state of the external world while the external operation continues or has already completed.
A single timeout can stand for at least three different states of the world, and each may require a different recovery strategy.
Idempotency, operation identity, postcondition verification, reconciliation, and explicit execution semantics are ways of managing that uncertainty. And as tool protocols increasingly expose behavioral metadata—MCP already provides hints for read-only, destructive, idempotent, and open-world behavior—the next engineering challenge is not merely collecting those declarations. It is understanding what evidence supports them and where declarations diverge from actual behavior.
That is the part worth investigating next—not another generic "AI agent reliability" framework.
The more precise question is whether we can make tool semantics observable, evidence-backed, and inspectable.