Numerical examples are illustrative unless explicitly sourced.
Durable execution preserves enough execution state and completed results for work to continue after a process failure. Its guarantees depend on the runtime, persistence configuration and the contracts of the systems it calls. It does not automatically make arbitrary external side effects occur exactly once.
Remember: Save progress. Repeat safely. Resolve uncertainty.
Understand the problem before learning the terminology
Imagine that a customer asks an agent to refund an order. The agent must read the order, check the policy, obtain approval, send the refund, and tell the customer. Each step takes time. The program can stop between any two instructions because a machine fails or a new version is deployed.
A worker is the running process doing that work. Its working memory is like notes on a whiteboard: useful while it is there, but not a reliable record after it disappears. Durable storage is a record that survives that particular worker, such as a database configured for the required durability. This does not mean the storage can never fail; it means the job's progress is not tied to one process's lifetime.
First separate two questions. Where did the job get to? is a progress question. Did the outside action already happen? is an outcome question. Saving progress helps answer the first. To answer the second, we may need cooperation from the payment service. Most confusion about durable execution comes from treating these as the same question.
Follow a complete run, then restart it
Suppose the job has ID job-42. The following is a conceptual history, not the exact event schema of a particular product:
| Recorded event | What the next worker can learn |
|---|---|
| Job started for order 42 | Which business request this run represents |
| Policy check completed: eligible | The completed check's stored result |
| Refund proposal created: $40 to original payment method | The specific action being considered |
| Manager approved proposal version 3 | Who approved which proposal |
Refund activity scheduled with operation ID refund-42-1 |
Which logical payment operation was requested |
Refund activity completed with receipt R901 |
A recorded successful result to reuse |
If the worker disappears after the policy check was recorded, a replacement can recover that result and continue to the approval stage. It need not ask a language model to reinterpret the policy merely to rediscover what the previous worker decided. If a policy change requires a fresh check, that is an explicit business rule, not an accidental consequence of a restart.
Replay means using the recorded events to rebuild the workflow's current state. Think of reconstructing a game's score from its recorded moves. You read the moves; you do not ask the players to play them again. Similarly, replay can supply a completed activity's stored result without sending the payment again. An activity whose completion was never recorded is a different case: it may need another execution attempt.
Deterministic orchestration means that, given the same recorded history, the workflow makes the same orchestration decisions. If replay reads today's wall-clock time directly and takes a new branch, its decisions can disagree with the old history. Replay-based engines provide supported ways to record time, randomness, and external results. The workflow says what happens next; activities perform work whose result may depend on the outside world.
The difficult case: money moved, but the receipt was not saved
Now stop the story one instruction earlier. The payment service sends the $40, but the worker crashes before recording receipt R901. The replacement worker sees a scheduled refund without a recorded completion. There are two possible realities: the payment never arrived, or it succeeded and the acknowledgment was lost. The local history looks the same in both.
An idempotency key solves this only when the receiving service enforces it. The agent resends refund-42-1; the service recognizes the same intended refund and returns its existing result rather than sending another $40. Creating a new key after each timeout defeats the protection. The service must also handle concurrent duplicate requests and define how long it remembers keys.
Reconciliation means comparing the uncertain local job with an authoritative external record and resolving the difference. For example, query the payment ledger by refund ID, recover the receipt, and mark the job complete. If the service cannot establish whether the refund happened, pause the job for investigation. “Unknown” is a legitimate state; treating it as “failed” can cost real money.
This is why neither “save done before the call” nor “save done after the call” is sufficient by itself. Saving before can leave a recorded success for money never sent. Saving after leaves a window where money was sent but success is absent. A transaction solves this within one database when it covers both operations. Across arbitrary services, you need their explicit protocols and contracts.
Waiting and changing the workflow
A durable approval wait stores that the job is awaiting proposal version 3. A signal or equivalent external event delivers the manager's decision to the job. A timer records when the wait expires. The workflow can release its worker while waiting and resume on another worker later. This avoids dedicating a sleeping process to every approval, but storage and the workflow service still consume resources.
A deployment introduces a separate problem. Yesterday's workflow may be waiting halfway through an old sequence. Today's code must still interpret that history correctly, or explicitly migrate it. Versioning means managing that compatibility. Test old histories before changing step order or action meaning. Correctly recovering yesterday's approval is not permission to execute today's changed proposal.
The five building blocks
| Term | Plain meaning | Refund example |
|---|---|---|
| Workflow | The rules for progressing through the job | Check → approve → refund → notify |
| Activity or task | A unit of work that may call the outside world | Call the payment API |
| Durable history/checkpoint | Progress stored beyond one process | Approval and completed activity result |
| Timer or external event | A recoverable wait or message | Wait for a manager until Friday |
| Idempotency key | Identity of one intended operation across retries | refund:tenant7:order42:request9 |
A key must represent the business operation, not a fresh random value on every attempt. A second legitimate partial refund needs a different operation ID. Reusing a key with changed parameters should fail, not silently issue a different refund.
Walk through the crash windows
Read diagram source
flowchart TD
A[Persist refund intent and stable operation ID] --> B[Check current permission and approval]
B --> C[Call payment service with operation ID]
C --> D{Outcome known?}
D -->|Success| E[Persist receipt and continue]
D -->|Confirmed rejection or failure| J[Record failure and choose permitted recovery]
D -->|Timeout or worker crash| F[Query status or retry same ID]
F --> G{Authoritative operation status}
G -->|Succeeded| E
G -->|Confirmed failure| J
G -->|Still unknown| H[Pause and reconcile with an operator]
| Failure point | What recovery knows | Safe behavior |
|---|---|---|
| Before the call | No attempt has started, if that boundary is reliably recorded | Execute under the normal policy |
| During the call | The request may have reached the receiver | Treat the result as unknown |
| After refund, before local result is saved | The money may already have moved | Reuse the same key or query the receiver |
| After result is durably recorded | The recorded activity finished | Reuse its result rather than issue a new refund |
An absent status record may mean the operation never arrived, is still in flight, is not yet visible or has aged out of retention. Treat absence according to the receiver's contract; it is not automatically a confirmed failure.
Exactly-once needs a stated boundary. A database can atomically insert a deduplication record and apply a local balance change. A remote service can offer a deduplicating API. A workflow journal alone cannot atomically commit both its own record and an arbitrary external effect. Key retention, concurrent attempts, and the receiver's contract all matter. Temporal explicitly documents the activity-completed/worker-crashed window and receiver-enforced keys. Temporal activity semantics.
If the receiver supports neither lookup nor deduplication, automatic retries of an ambiguous write may be unsafe. Pause for reconciliation or redesign the integration. Marking the local step “done” before calling merely exchanges duplicate risk for lost-action risk.
Replay, checkpoints, and model calls
Replay reconstructs workflow state from recorded events. Completed recorded activities can return stored results; unfinished work may execute again. Therefore, record model responses as activity results when replay must preserve the original decision. An unrecorded model call can run and be billed again.
In a replay-based engine, orchestration must follow its determinism rules. Use supported clock/randomness primitives or activities, not arbitrary network calls in replayed orchestration. Checkpoint systems save graph state at defined boundaries; durable storage and the chosen persistence mode determine what survives a crash. In-memory checkpointers do not survive process loss. These are implementation contracts, not a simple “checkpoints bad, workflows good” divide. LangGraph persistence.
A persisted wait does not require a dedicated sleeping worker, but the service and storage still have cost. Recovery restores execution state; it does not restore the outside world or secretly reveal the model's internal thoughts.
Approval, cancellation, and compensation
Bind approval to the exact action, amount, recipient, and version of the proposal. Persist who approved it and its expiry. On resume, recheck permissions and relevant business state. Changing the amount invalidates the old approval.
Cancellation is a request to stop further work, not proof that an in-flight write was prevented. If the refund committed, cancelling the workflow does not undo it. A compensating action is a new business operation, such as cancelling a reservation; it can itself fail and may not restore the original situation.
Choose the smallest sufficient design
| Situation | Reasonable starting point | What still needs engineering |
|---|---|---|
| Cheap read-only job that can restart | Queue, bounded retries, result record | Deduplication, deadlines, dead-letter handling |
| Conversation/graph with resumable steps | Persistent graph checkpointer | Durable backend, replay boundaries, safe tools |
| Hours of work, approvals, many services | Durable workflow engine | Activity contracts, history growth, versioning, operations |
| Updates contained in one database | Transaction plus durable job table | Atomicity and concurrency within that database |
Temporal, Restate, DBOS, Inngest, and Step Functions are options to investigate, not a universal ranking by footprint. Compare the actual deployment, supported languages, persistence guarantees, timers, concurrency control, version migration, and your team's ability to operate it.
What an AI manager should own
Assign one owner for workflow recovery and one for each side-effect contract. Define a deadline for stuck approvals, a manual reconciliation queue, and metrics for recovery time, unknown outcomes, duplicate effects, retries, and cost per completed job. Test crashes before and after the receiver commits, duplicate approval events, worker restarts, and permission revocation while paused.
Before selecting a runtime, write down:
- The maximum acceptable lost progress and recovery time.
- Which operations are safe to repeat and which need receiver deduplication.
- How long operation IDs, approvals and result records remain valid.
- How concurrent workers are fenced from conflicting updates.
- How pending work survives code/schema changes.
- Who resolves unknown outcomes and how the queue is monitored.
Persist compact records and references instead of embedding every large document in each checkpoint. For illustration, 100,000 runs/day × 40 records/run × 2 KB/record produces 8 GB/day, or 240 GB over 30 days, before replicas and indexes. Repeating a 1 MB artifact in all 40 records instead would produce 4 TB/day. Actual engines have different event formats and storage behavior; measure them and use their supported history/continuation mechanisms.
Deploy changes compatibly with running histories: use the engine's versioning/migration mechanisms and test replay against representative old runs. A new deployment must not silently reinterpret yesterday's approved action.
Recall and follow-up questions
“Does Temporal guarantee exactly-once payments?” No. Durable orchestration plus the payment service's deduplication/transaction contract can provide that effect within a defined scope. The journal alone cannot.
“Why not just retry?” A timeout tells me I did not receive success; it does not tell me the action failed.
“Why use this for a read-only agent?” A six-hour analysis may justify it to preserve expensive progress even without irreversible actions.
“How would you prove recovery works?” Kill workers at the crash boundaries and verify the actual payment ledger and final job state, not just the agent's message.
Close the page and draw the refund failure window. If you can explain why “save before” and “save after” each leave a problem, you understand the core concept.
Related: Recovery, Human approval, Reliability.
Choose a runtime by its recovery boundary
The comparison below focuses on where each system records progress and controls concurrent work. Check the exact SDK, persistence configuration, and deployment model. All choices still need application authorization and a safe contract for external writes.
| Option | Useful mental model | Boundary to explain in an interview |
|---|---|---|
| Temporal | Durable workflow history plus retriable activities and messages | Replay-compatible orchestration and activity side effects need different treatment; message acknowledgment is not always business completion |
| Restate | Durable handlers, journaled operations, and keyed services | Exclusive handlers on a virtual-object key serialize state mutation; shared handlers and different keys have different concurrency semantics |
| DBOS | Workflows with recorded steps and tracked database transactions | A supported datasource transaction records its outcome atomically with database effects; an arbitrary HTTP payment is outside that atomic boundary |
| Inngest | Event-triggered functions broken into persisted, retriable steps | Put appropriate work inside steps and reason about retried external calls; persisted step results do not deduplicate an uncooperative receiver |
| AWS Step Functions | Managed state machines integrating services | Standard and Express have different duration, history, and execution semantics; a task's configured retries can still repeat its external effect |
| LangGraph | Stateful graph execution with checkpoints and resumable tasks | Durability depends on the checkpointer and persistence mode; task boundaries must isolate nondeterminism and effects appropriately |
For example, a Restate virtual object keyed by tenant/order can own the order's current refund proposal and serialize conflicting updates. This resembles a per-order controller: messages for other orders can proceed independently. A shared read handler does not acquire the same exclusive mutation behavior. Do not put every customer under one global key unless global serialization is intended.
Signal, query, or update?
In Temporal, a Signal asynchronously delivers a message that can change workflow state; acceptance does not mean the handler has completed the business operation. A Query reads workflow state without adding an event to history and cannot mutate it or block waiting for work. An Update provides a tracked request/response interaction that can validate and mutate state and return a result. Use the message-passing contract, not “all messages are signals.”
For our refund, “manager approved proposal 3” is a state-changing message. “What is its status?” is a query. “Validate and register this approval, and tell me whether it was accepted” may fit an update. None grants permission to change proposal 3 into a different amount. Handlers that yield can interleave under the engine's execution model, so protect application invariants.
Sync, async, and exit checkpoint persistence
LangGraph documents three durability modes. Sync persists a checkpoint before the next step; async persists while subsequent work proceeds; exit persists when execution exits rather than at every intermediate step. Async reduces waiting but introduces a crash window for a pending write. Exit reduces intermediate persistence overhead but can lose more progress on process failure. A configured durable backend is still necessary. See the current durability type reference.
Practice the difference with the same timeline: node A produces a model answer, node B proposes a refund, and the process dies during B. Ask which A result was durably stored under the chosen mode, what will be recomputed, and whether any external action needs reconciliation. Choosing sync narrows a state-loss window; it cannot make a separate payment and checkpoint one transaction.
Recall sentence: Pick the engine by the state, messages, concurrency, and recovery contract you need; pick the side-effect protocol separately.
Interview questions with developed answers
Q1: Why are naive retries and checkpoints insufficient for an agent with side effects?
Sample answer: A checkpoint tells me the latest state the application saved. It cannot, by itself, settle a payment that succeeded immediately before a crash. Retrying that payment with a new identity can pay twice. I use durable workflow history to preserve progress and completed results, and a stable operation ID enforced by the receiving payment service to make repeated attempts safe. If the receiver provides a status lookup, I can recover the existing result. If neither deduplication nor lookup exists, I must pause and reconcile the uncertain write. A workflow engine manages execution; it does not magically make every external action atomic.
Why this works: It separates remembering progress from preventing duplicate business effects.
Follow-up — What if the key has expired? The deduplication guarantee may no longer apply. Check the receiver's retention contract and authoritative ledger before resubmitting an old operation.
Q2: Explain replay without saying that it reruns the entire agent.
Sample answer: Replay rebuilds orchestration state from recorded events. If the policy-check activity completed and its result was recorded, the workflow can consume that stored result while rebuilding state. It does not need a new policy-check model call just because a worker restarted. The same is true of recorded tool results. But an activity whose completion was not recorded may be attempted again. I put network calls and other nondeterministic work behind the engine's supported activity or task boundaries and design those attempts safely. That preserves past decisions while allowing unfinished work to progress.
Why this works: It distinguishes re-executing orchestration code from repeating external actions.
Follow-up — Is an LLM deterministic at temperature zero? Do not rely on that for recovery. Persist the actual response when the workflow must preserve it.
Q3: When is durable execution overkill, and what would you use instead?
Sample answer: A short, inexpensive summary that can safely restart may only need a queue, bounded retries, and a result record. A persistent graph checkpointer may be sufficient when the main need is resuming conversation or graph state. I consider a workflow engine when jobs wait a long time, coordinate many services, or lose substantial work on restart. Side effects increase the need for careful contracts, but a six-hour read-only analysis can also justify durability. I compare recovery requirements and operational burden, rather than choosing a product because it appears on an agent architecture diagram.
Why this works: It weighs the cost of lost progress against the complexity of recovery. Expensive read-only work can justify durability too.
Follow-up — What changes your decision? Measured restart cost, growing approval waits, and the complexity of custom recovery code.
Q4: How do approval and cancellation behave after a restart?
Sample answer: I persist the exact proposal, its version, approver, and expiry. The resume event references that proposal. Before execution I check that permissions and relevant business conditions still allow it. A changed amount needs new approval. Cancellation stops future work where possible, but an in-flight refund may already have committed. I first establish its outcome; any compensation is a separate business action with its own permission and failure handling. Restoring an old checkpoint does not reverse a real payment.
Why this works: It explains both authorization and the limits of recovery.
Follow-up — What about duplicate approval messages? Deduplicate them and allow only valid state transitions, so one approval cannot initiate two independent refunds.
Q5: How would you test this design and choose a workflow tool?
Sample answer: I kill workers before the call, during the call, after the receiver commits, and after the local completion record. I inspect the actual ledger, not just the agent's final message. I also test duplicate events, expired approvals, receiver outages, and a deployment resuming an older workflow. For tools, I compare persistence boundaries, timers, concurrency controls, language support, versioning, and operating responsibility. Temporal, database-backed workflow libraries, managed state machines, and graph persistence solve overlapping needs with different contracts. I would prototype the hardest failure window before committing the team.
Why this works: It turns a conceptual guarantee into observable evidence and an adoption decision.
Follow-up — Who owns stuck runs? A named operational owner needs a reconciliation queue, alerts, a runbook, and authority to resolve or safely terminate them.
60-second interview answer
Durable execution lets a job continue after the worker running it disappears. I would persist progress, model results, approvals, and timers so recovery does not start the whole agent again. But remembering progress does not itself prevent duplicate payments or emails. A tool can succeed just before the worker crashes, leaving its outcome unknown. I need the receiving service to deduplicate retries using a stable operation ID, or I must reconcile its state before proceeding. I would adopt a workflow engine when long waits, expensive work, or complex recovery justify its operational cost.