What a trading engine taught me about agent retries
Exchange infrastructure solved idempotency, retry storms, and partial failure decades ago. Agent frameworks are rediscovering all three the hard way.
There is a moment in every agent project where someone adds a retry. The tool call timed out, the run failed, and a retry makes the demo work again. It ships.
Six weeks later the same retry is why a customer got charged twice.
The problem is not retrying. It is retrying without identity.
On the trading side we had a rule that predates every framework in your
package.json: a message that can be retried must carry an identity that makes
the retry a no-op. Not a timestamp. Not a hash of the payload. A caller-supplied
key that survives the network, the queue, and the operator who reruns the batch
at 2am because a dashboard was yellow.
Agent frameworks default to the opposite. The model decides to call
create_invoice, the call times out, the framework retries, and the tool has no
way to know these two requests are the same intent. You get two invoices.
What to do instead
Give every tool invocation an idempotency key generated at the call site, before the first attempt, and make the tool contract require it:
import uuid
def invoke_tool(name: str, args: dict, *, attempt_key: str | None = None) -> dict:
# Generated once per logical intent, reused across every retry of that intent.
key = attempt_key or str(uuid.uuid4())
return tool_registry[name](**args, idempotency_key=key)
The key is generated once per intent, not once per attempt. Reuse it for the whole retry sequence. On the receiving side, store it and return the original result on a repeat.
Retry storms are worse than the original failure
The second thing exchange systems taught, expensively: when a dependency degrades, every client retrying in lockstep converts a slow dependency into a dead one. Exponential backoff alone does not fix this — synchronized clients back off in synchronized waves.
You need jitter, and you need a budget:
import random, time
def with_retries(fn, attempts=3, base=0.4, cap=8.0, budget_s=20.0):
deadline = time.monotonic() + budget_s
for i in range(attempts):
if time.monotonic() > deadline:
raise TimeoutError("retry budget exhausted")
try:
return fn()
except TransientError:
if i == attempts - 1:
raise
delay = min(cap, base * 2 ** i)
time.sleep(random.uniform(0, delay)) # full jitter, not fixed backoff
The budget matters more than the attempt count. An agent with a twelve-step plan and three retries per step has a worst case of thirty-six calls and a latency profile nobody modeled. A wall-clock budget bounds the damage.
The instinct that transfers
Ask, for every tool your agent can reach: what happens if this runs twice? If the answer is "nothing," you are fine. If the answer is "the customer notices," you have an idempotency requirement, and the model is not going to enforce it for you.
That question is thirty years old. It did not stop being the right question when we started letting a language model choose the call order.
Subscribe
The weekly essay, free.
One piece a week on what holds up in production. No digest, no roundup, no course funnel.
Related
Determinism is a budget, not a property
You cannot make an LLM deterministic. You can decide exactly how much nondeterminism each part of your system is allowed to spend, and enforce it at the boundaries.