Make Long-Running Agents Resume from Durable State

Design agent jobs around persisted decisions and accepted effects, so an interruption does not force the system to guess what already happened.

A long-running agent can stop after the model has made a decision, after a tool request has been sent, or after an external service has accepted the request. Those are different recovery situations. Replaying the conversation from the beginning does not explain which one occurred.

I design resumable work around durable application state. The conversation can help explain intent, but recovery needs explicit records of the job, the completed steps and the effects whose outcomes are still uncertain.

Give the job an identity beyond the session

A job should have a stable identifier that survives worker restarts and browser disconnections. That identifier connects its requested objective, account scope, inputs, progress and output. A new session can then inspect the existing job instead of creating unrelated work by accident.

I would store the versions of inputs that influenced a consequential decision. If the source data changes while the job is paused, resumption may need a fresh decision. A saved plan is useful context; it is not proof that the original assumptions still hold.

This is a recurring architectural concern in the operational systems work I undertake through Wizutech. Business processes can outlast a single request, and recovery should preserve their identity.

Checkpoint meaningful boundaries

A checkpoint is most useful when it captures a state that the application knows how to resume. For example, a research job might distinguish sources collected, findings reviewed and publication proposed. Those states describe business progress rather than a position inside generated text.

The durable record should contain accepted outputs and unresolved questions. I would avoid making recovery depend on reconstructing private model reasoning. The application needs the decision, relevant evidence and execution state, not a transcript of every intermediate thought.

Each resumption also needs compatible code and data. If a new release changes the shape of saved state, the worker should migrate or reject that version explicitly instead of interpreting old fields under new assumptions.

Handle the gap around external effects

The difficult case is a tool call whose remote effect succeeded before the local worker recorded success. Retrying blindly can repeat the effect. Marking the step complete before making the call creates the opposite failure: the job can claim success for an operation that never happened.

Where available, I use a stable operation key and a way to look up the remote outcome. Otherwise the job should preserve an uncertain state and take a reconciliation path. Persistence alone cannot create an exactly-once guarantee across an unrelated external service.

The design connects directly to separating proposed actions from accepted effects. A model’s intention and a tool’s confirmed result belong in different fields.

Prevent two workers from resuming the same step

A restart can overlap with a slow original worker. A lease or ownership record helps coordinate execution, but it needs a way to reject work from an expired owner. I would check the current execution version at the write boundary, not assume that an earlier lease check remains valid forever.

For external operations, local coordination still needs the destination’s duplicate protection or reconciliation behavior. A database lock cannot retroactively cancel a request already accepted elsewhere.

Make recovery visible

The operator should be able to tell whether a job is running, waiting for a dependency, reconciling an uncertain effect or ready to resume. “Retrying” is too vague if the system is actually waiting for a person to clarify the original request.

I would review each interruption point alongside the job’s stop conditions. Durable state is valuable when it allows the next worker to continue a known process. It should reduce the amount of inference required after a failure, especially around actions that affect other systems.

Updated 30 September 2026.