Skip to content
Load-Bearing CodeSubscribe
August 20, 2026529 words2 min

What a trading engine taught me about agent retries

Exchange infrastructure solved idempotency, retry storms, and partial failure decades ago. Agent frameworks are rediscovering all three the hard way.


There is a moment in every agent project where someone adds a retry. The tool call timed out, the run failed, and a retry makes the demo work again. It ships.

Six weeks later the same retry is why a customer got charged twice.

The problem is not retrying. It is retrying without identity.

On the trading side we had a rule that predates every framework in your package.json: a message that can be retried must carry an identity that makes the retry a no-op. Not a timestamp. Not a hash of the payload. A caller-supplied key that survives the network, the queue, and the operator who reruns the batch at 2am because a dashboard was yellow.

Agent frameworks default to the opposite. The model decides to call create_invoice, the call times out, the framework retries, and the tool has no way to know these two requests are the same intent. You get two invoices.

What to do instead

Give every tool invocation an idempotency key generated at the call site, before the first attempt, and make the tool contract require it:

import uuid

def invoke_tool(name: str, args: dict, *, attempt_key: str | None = None) -> dict:
    # Generated once per logical intent, reused across every retry of that intent.
    key = attempt_key or str(uuid.uuid4())
    return tool_registry[name](**args, idempotency_key=key)

The key is generated once per intent, not once per attempt. Reuse it for the whole retry sequence. On the receiving side, store it and return the original result on a repeat.

Retry storms are worse than the original failure

The second thing exchange systems taught, expensively: when a dependency degrades, every client retrying in lockstep converts a slow dependency into a dead one. Exponential backoff alone does not fix this — synchronized clients back off in synchronized waves.

You need jitter, and you need a budget:

import random, time

def with_retries(fn, attempts=3, base=0.4, cap=8.0, budget_s=20.0):
    deadline = time.monotonic() + budget_s
    for i in range(attempts):
        if time.monotonic() > deadline:
            raise TimeoutError("retry budget exhausted")
        try:
            return fn()
        except TransientError:
            if i == attempts - 1:
                raise
            delay = min(cap, base * 2 ** i)
            time.sleep(random.uniform(0, delay))   # full jitter, not fixed backoff

The budget matters more than the attempt count. An agent with a twelve-step plan and three retries per step has a worst case of thirty-six calls and a latency profile nobody modeled. A wall-clock budget bounds the damage.

The instinct that transfers

Ask, for every tool your agent can reach: what happens if this runs twice? If the answer is "nothing," you are fine. If the answer is "the customer notices," you have an idempotency requirement, and the model is not going to enforce it for you.

That question is thirty years old. It did not stop being the right question when we started letting a language model choose the call order.

Subscribe

The weekly essay, free.

One piece a week on what holds up in production. No digest, no roundup, no course funnel.

← All writing