Durable Agents: When Step Seven Fails
A demo agent runs for ten seconds inside one HTTP request. A production agent runs for minutes, sometimes hours, calls a dozen tools, and writes to systems that matter. Somewhere in there something fails. A tool times out. A provider returns a 500. A pod is evicted mid-deploy. What happens next is the line between a prototype and a system.
The obvious answer is to retry the run. That is also the answer that creates duplicate invoices.
Agents belong on a queue
Most teams get the first decision wrong: where the agent runs. An agent running inside the HTTP request that triggered it has its lifetime coupled to a connection, a load balancer timeout, and a deploy cycle. None of those should get a vote on when your agent stops.
Put the run on a durable queue. The API accepts the request, enqueues it, and returns a run id immediately. A worker picks it up. The client polls or subscribes for progress. Now the run survives a deploy, and workers scale independently of the API.
Checkpoint the trajectory
Durable execution means the run is a state machine and each transition is persisted before the next one begins. For an agent, the state worth persisting is the trajectory: the sequence of steps, tool calls, arguments, results, and accumulated context. Not the model's internals, which aren't yours to checkpoint anyway.
type RunState = { runId: string; status: "running" | "waiting_approval" | "failed" | "done"; step: number; trajectory: Step[]; // append-only pendingCall?: ToolCall; // written BEFORE the call executes};The detail that matters is pendingCall. You persist the intent to call a tool before calling it. If the process dies mid-call, recovery knows a call was in flight and can decide what to do, rather than guessing whether step seven happened.
Here's what gets checkpointed.
Idempotency is what makes retry safe
Every side-effecting tool call carries an idempotency key derived from the run id and the step number. The downstream system uses it to deduplicate. Retry becomes safe by construction. A common version of the failure this prevents: step seven is the invoice call, the process dies after the request leaves but before the result is recorded, and the retry sends it again. Without the key the customer gets two invoices. With it, the second call returns the first result.
await tools.createInvoice(args, { idempotencyKey: `${state.runId}:${state.step}`,});If a tool you depend on has no idempotency support, you own the deduplication. Record the call and its result under that same key before marking the step complete, and check the record on resume.
Some steps can't be undone
Idempotency stops you doing something twice. It doesn't help when the correct response to a failure is to undo what already happened.
For those you need one of two things. A compensating action, defined alongside the tool that requires it, so the run can walk backward. Or a human approval gate before the irreversible step, so nothing unrecoverable happens without a person confirming it. Most systems need both, applied to different classes of action.
If an irreversible action has neither a compensation nor an approval gate, that's a design gap.
An approval gate also means the run parks, possibly for hours or overnight. That's a timer, and it's exactly what durable execution engines are good at. A worker sleeping on a thread while it waits for a human is a worker you are paying for and an outage waiting for the next deploy.
On earning the right to act without approval: Evals for Production Agents.
Resume or restart
On recovery you choose. Resume from the last checkpoint when the trajectory is still valid and the world hasn't moved underneath you. Restart when too much time has passed, when the underlying data may have changed, or when the failure suggests the plan itself was wrong. A restart gets a new run id, so its idempotency keys can't collide with the old plan's steps.
Make this a policy per run type. A research run that gathers information can safely restart. A run that has already moved money can't.
What to log
The full trajectory in a form you can replay. Every tool call with arguments, idempotency key, result, latency, and cost. Every checkpoint write. Every resume, with the reason.
When an agent does something surprising three weeks from now, this log is the only thing that will explain why. It also happens to be exactly the data your trajectory evals need. Chain of Thought Is Not an Audit Trail covers what that log has to contain.
Do not build the state machine yourself
Durable execution is a solved problem with production tooling behind it. LangGraph ships checkpointers that persist graph state between steps, which covers the agent-shaped case with the least ceremony. Temporal, Restate, and DBOS solve the general problem, with retries, durable timers, and compensation as first-class concepts, and they are the right answer when a run spans systems and model calls.
Write your own only when a hard constraint rules those out. The failure modes here are subtle, they surface under load, and they are expensive to learn firsthand.
Working on something like this?
Start a Conversation