System designby Learnastra

Concept lesson · Foundations

Idempotency, retries, and timeouts

By Anup Rai

Start here

Definition

An operation is idempotent when repeating the same logical request has the same intended effect as performing it once. A retry is another attempt at that request; a timeout only says the caller stopped waiting and does not establish whether the effect happened.

Why it matters: Networks can lose the response after a server commits. A client needs a way to recover the original result without accidentally creating another order or charge.

The visual modelIdempotency keys and recovery after a lost response

The server binds the caller and idempotency key to the request fingerprint and committed result. A retry must not create a second business operation.

Idempotency keys and recovery after a lost responseThe server binds the caller and idempotency key to the request fingerprint and committed result. A retry must not create a second business operation. An authenticated client U9 sends POST /orders with idempotency key buy-204. The server commits order O17 and the result for (U9,buy-204) together. The reply is lost. Retrying the same request and identity returns the stored O17 result. Reusing buy-204 with a changed payload is rejected. A distinct intended purchase uses a new key. Bound retries with backoff, jitter and an end-to-end deadline; retry safety is not overload control.Lost reply after commit: recover O17CLIENT / U9DURABLE ORDER AUTHORITY1. buy-204; item B2; quantity 1; quote Q8Commit O17 + request result together2. response lost3. same U9 + buy-204 + payload4. return O17; no second orderSame identity + changed payload = conflict. The unique request result guards the commit.
Read the diagram step by step
  1. An authenticated client U9 sends POST /orders with idempotency key buy-204. The server commits order O17 and the result for (U9,buy-204) together.
  2. The reply is lost. Retrying the same request and identity returns the stored O17 result.
  3. Reusing buy-204 with a changed payload is rejected. A distinct intended purchase uses a new key.
  4. Bound retries with backoff, jitter and an end-to-end deadline; retry safety is not overload control.

Worked example

U9 submits buy-204 and the server commits order O17, but the reply is lost. A retry with the same caller, key, and payload returns O17 instead of creating O18.

Key takeaways

  • A timeout means unknown outcome, not proven failure.
  • Commit the request identity, the local business change, and its saved result together so a crash cannot separate them.
  • Bound retry attempts and elapsed time; external effects need their own recovery contract.

You will learn to

  • Distinguish a failed operation from an unknown outcome.
  • Design an atomic idempotency record with a request fingerprint.
  • Budget retries and isolate overloaded dependencies.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: HTTP APIs and request lifecycle · Transaction isolation

Workload and timing examples are interview assumptions.

01Idempotency, retry, timeout, and deadline: definitions

Idempotency means that repeating the same logical operation has the same intended business effect as performing it once. A retry is another attempt at that operation. A timeout is a maximum waiting duration: if no response arrives, the caller does not know whether the server received the request, committed it, or lost its reply. A deadline is an absolute point in time by which a call should finish. Propagating one deadline bounds the total waiting budget across a call chain; it does not prove that remote work stopped or undo committed effects.

Idempotency does not require every low-level network message to occur once. Without a stable operation identity, another attempt can accidentally create a second business intent. Our target is one order O17 for checkout attempt buy-204, even if three HTTP attempts arrive.

This distinction appears in payments, file uploads, job queues, webhooks, and agent tool calls. First identify the business operation and which transaction or external service durably records its result. Then decide how a retry finds that result.

02Lost-response retry: one order from two attempts

Bind an idempotency key, a stable identifier for one logical operation, to the authenticated caller and a normalized representation of the request. For example, POST /orders uses Idempotency-Key: buy-204, user U9, item B2, quantity 1, and quote Q8. The following trace isolates the failure between committing O17 and returning its response.

Step Server state Caller-visible state
1 No record for (U9, buy-204) Request sent
2 Transaction creates O17 and saves the request result Still waiting
3 Transaction commits Still waiting
4 Response is lost Timeout; outcome unknown
5 Same key arrives again Same logical operation retried
6 Server returns saved O17 result Purchase confirmed once

The request identity must come from a stable retryable intent, not a fresh random key on every network attempt. A distinct second purchase should use a new key. Reusing a key with a different item should be rejected rather than silently returning a result for the wrong request.

Worked example diagramThe request-result record protects the local order identity. The payment remains a separate effect with its own retry and reconciliation contract.
Idempotency, retries, and timeouts: architecture diagram1. U9: buy-204 to 2. Order API: first attempt or retry; 2. Order API to 3. Unique request-result record: claim caller/key atomically; 3. Unique request-result record to 4. Order O17: commit order and result together; 4. Order O17 to 5. Payment attempt pay-204: durable external attempt identity; 3. Unique request-result record to 6. Retry: return O17: retrieve committed outcome1 → 2: first attempt or retry2 → 3: claim caller/key atomically3 → 4: commit order and result together4 → 5: durable external attempt identity3 → 6: retrieve committed outcome01U9: buy-20402Order API03Uniquerequest-resultrecord04Order O1705Payment attemptpay-20406Retry: return O17
  1. 1 → 2first attempt or retryU9: buy-204 → Order API
  2. 2 → 3claim caller/key atomicallyOrder API → Unique request-result record
  3. 3 → 4commit order and result togetherUnique request-result record → Order O17
  4. 4 → 5durable external attempt identityOrder O17 → Payment attempt pay-204
  5. 3 → 6retrieve committed outcomeUnique request-result record → Retry: return O17

03Idempotency key, payload fingerprint, and atomic result storage

Store RequestResult(callerId, key, payloadHash, state, resourceId, response) with a unique (callerId, key) constraint. A payload hash is a fingerprint computed from the fields that define the operation, such as item, quantity, and quote. Canonical means these fields are normalized consistently before hashing, so equivalent inputs produce the same representation. It detects reuse of the same key for a different intent; it is not authorization.

When work cannot finish in one short database transaction, persist its progress and give a worker temporary ownership, often through a lease. Expiry lets a replacement take over, but recovery must still account for requests the previous worker may already have sent. This is why a long operation needs more states than simply “key absent” or “completed.”

Specify the key namespace: the group within which an idempotency key must be unique, such as all requests by one caller. The example uses caller-wide keys, so the fingerprint includes the operation and target as well as item fields. A tenant or service that uses separate namespaces must include that scope in the unique identity. Replaying a saved response still requires current permission; an old idempotency key must not expose a resource after access is revoked.

Stored state Same identity and payload Unsafe reaction
No record Atomically create the effect and outcome, or durably claim a long operation Check absence and create outside one protected boundary
In progress Return status, wait within budget, or recover ownership Launch another uncoordinated worker
Completed Return the recorded effect identity and an authorized result Repeat the business mutation
External outcome unknown Reconcile the original external operation Treat timeout as rejection and choose a fresh key
Same key, different fingerprint Reject the conflict Return an unrelated old result

A lease lets a replacement worker take over after a deadline. The old worker may resume later, so the store must atomically check the current ownership version (epoch) and expected state before saving a result. That check cannot undo an external request already sent. The receiving service still needs duplicate protection, or a way to check and resolve the uncertain result.

04External effects and uncertain payment outcomes

Suppose checkout calls a payment provider after creating an order. The provider charges successfully, but its reply is lost before local state records success. Repeating a new provider request can double-charge even if the local order insert was idempotent.

Use one stable provider attempt key for the payment, record it durably before or as part of scheduling the attempt, and reconcile the provider's status after uncertainty. A webhook may report the result, but duplicate and reordered webhooks need their own identity/state checks. Only finalize the local purchase once the confirmed result satisfies the state machine.

If the provider has no safe retry or status lookup, an uncertain payment may need manual investigation. Explain that limit. To claim duplicate protection, identify the exact action protected, how long its request ID is remembered and which failures are covered. “Exactly once” alone explains none of those.

Read the provider's actual contract rather than copying a generic retry recipe. For example, Stripe documents replaying the first saved status and body for an idempotency key, including a saved 500. Reusing that key can therefore replay an error without proving that no effect occurred; using a fresh key simply to escape the saved error can duplicate work. Keep the operation pending and use the supported recovery path.

05End-to-end deadlines and retry amplification

A deadline is an absolute point in time by which a call should finish; a timeout is a maximum waiting duration, often for one step. If the user allows two seconds for checkout, giving three nested services independent two-second timeouts can exceed that budget. Propagate the deadline or its remaining time budget through the call chain and stop work that is no longer useful when safe to do so.

Concept in focusOne request can become 27 storage attempts

Every parent branches into three total attempts, including the original. Read from top to bottom.

One request can become 27 storage attemptsEvery parent branches into three total attempts, including the original. Read from top to bottom. Count the branching levels: 1, 3, 9, 27. Three caller attempts each permit three middle-layer attempts. Each of those nine can permit three storage attempts, producing 27 in the worst case.Independent retries multiply: 3 x 3 x 3 = 27 attempts1 logical request3 caller attempts9 middle-layer attempts27 storage attemptsCount includes the first attempt. Retry at selected layers within one deadline.

Remember: Retry budgets multiply across layers.

Read the diagram
  1. Count the branching levels: 1, 3, 9, 27.
  2. Three caller attempts each permit three middle-layer attempts.
  3. Each of those nine can permit three storage attempts, producing 27 in the worst case.
Try from memoryIf only the outer layer permits three attempts, how many storage attempts can one request cause?

At most three in this simplified chain, assuming each inner layer makes one attempt per call.

Retry transient transport failures or documented retryable responses when the operation is safe and time remains. Do not repeatedly retry invalid input, denied permission, or a business condition that is no longer satisfied, such as an expired reservation. Respect server retry guidance.

Budget connection setup, queueing, processing, backoff and response transfer within the same end-to-end limit. If Retry-After asks for a wait beyond the remaining interactive budget, return a retryable/pending result instead of sleeping and then starting an already-expired attempt. HTTP and RPC clients may have their own automatic retries, so inventory them before multiplying attempts. gRPC clients also need an explicit realistic deadline; deadline propagation and cancellation handling vary by language and application code.

06Exponential backoff, jitter, circuit breakers, and bulkheads

Bounding the number of retries still leaves two problems: many clients may retry together, and slow calls may occupy every available resource. The controls below address different parts of that load: when to retry, whether to call a failing dependency, and which workloads share a resource pool.

Exponential backoff increases the waiting interval between retries. For a base interval of 100 ms, caps might be 100, 200, 400, and 800 ms. Jitter randomizes each wait, for example choosing a value between zero and the current cap. Ten thousand clients then avoid retrying at exactly the same instant.

A circuit breaker stops calls temporarily after sufficient failure evidence and later allows limited probes. It reduces repeated futile work; it does not repair the dependency or authorize dropping important writes. A bulkhead gives workloads separate concurrency/resource pools so a slow image export cannot consume every checkout connection.

Control What it bounds Example
Deadline Total useful elapsed time Stop interactive checkout attempts after its budget
Retry budget Additional attempts At most one extra request at the API layer
Backoff + jitter Timing of repeated work Spread recovery attempts
Concurrency limit Work in flight Only 50 simultaneous provider calls
Circuit breaker Calls into known failure Limited recovery probes

Queues also need limits. If work arrives faster than it can complete indefinitely, an ever-growing queue delays the failure while consuming memory or storage; it does not add processing capacity.

07Interview walkthrough: safe checkout retries

Interviewer: “The customer presses Buy twice because the first request timed out. How do you prevent two orders?”

Candidate: “Both attempts use buy-204 for the same user and request. I save the unique request-result record and order in one transaction. If the reply is lost after commit, the retry returns O17. For an external payment, I reuse the provider’s request key and check uncertain results. I stop retries at the user’s deadline and spread them with backoff and jitter so an outage does not trigger a flood.”

The answer is grounded because it names the durable state before and after the lost response, rather than assuming the network delivers exactly once.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What does idempotency mean, and why does a timeout make it useful?

Reveal a model answer

Idempotency means repeating the same logical operation has the same intended effect as doing it once. A timeout does not prove failure: O17 may have committed before its response was lost. U9 retries buy-204 with the same caller and payload so the service returns O17 instead of making another purchase.

What the answer must demonstrate: Distinguish one business effect from one transport attempt.

Foundation · Question 2

What exactly does the idempotency key identify?

Reveal a model answer

“One caller’s logical intent, such as U9’s purchase attempt buy-204. I scope it to the authenticated caller and compare a canonical payload fingerprint. A fresh network retry reuses it; a new intended purchase gets a different key.”

What the answer must demonstrate: Separate caller, intent, and payload.

Applied · Question 3

Two requests both see no saved result. How is one order guaranteed?

Reveal a model answer

“The claim and business effect must share an atomic boundary, such as a unique request record and order insert in one transaction. A separate check-then-insert can let both proceed. The losing concurrent attempt waits for or retrieves the winner’s outcome.”

What the answer must demonstrate: Show the atomic boundary.

Applied · Question 4

Can idempotency records expire after a minute?

Reveal a model answer

“Only if the contract prevents valid retries after that minute or another durable identity prevents repetition. Deleting the record can make a delayed duplicate look like a new operation. I align retention with the retry horizon, business identifiers, and downstream retention.”

What the answer must demonstrate: Treat deduplication retention as part of correctness.

Applied · Question 5

Why doesn’t a local transaction make the external charge exactly once?

Reveal a model answer

“The provider is outside that transaction. It can charge and lose its reply before we save the result. I use one durable provider attempt key, query or reconcile its outcome, and process duplicate notifications safely. The local order and remote charge have separate commit boundaries.”

What the answer must demonstrate: Avoid blanket exactly-once claims.

Applied · Question 6

Three layers each make three attempts. What reaches the bottom?

Reveal a model answer

“In the worst simple nesting, up to 27 calls for one user operation. That extra work can keep a struggling dependency down. I choose one retry layer or a shared budget, cap attempts and total time, and stop retrying permanent failures.”

What the answer must demonstrate: Show the multiplication and the bound.

Applied · Question 7

How do timeouts relate to a two-second user budget?

Reveal a model answer

“The two-second budget becomes an absolute deadline two seconds after the request starts. Each downstream call gets at most the remaining time as its timeout, including planned retries. Independent two-second waits at every layer can greatly exceed it and keep doing work after the user has left.”

What the answer must demonstrate: Distinguish stopping work from reversing it.

Applied · Question 8

How do you stop a slow export dependency from taking down checkout?

Reveal a model answer

“I separate concurrency pools so exports cannot consume all checkout workers or connections. I bound queues and use deadlines, then reject or defer lower-priority work when capacity is exhausted. A breaker can limit calls to the unhealthy dependency while probing recovery.”

What the answer must demonstrate: Protect a finite resource and explain overload behavior.

Blank-page exercise · 20 minutes

Build the answer yourself

Draw a purchase that commits before its response is lost. Add a concurrent retry and an uncertain payment outcome.

  • Name the request key, caller, and payload fingerprint.
  • Show the transaction and external-effect boundaries.
  • Specify duplicate handling and retention.
  • Calculate retry amplification and set a bounded budget.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Idempotency, retries, and timeoutsTimeoutRecall first, then reveal

The caller stopped waiting; the durable outcome may already exist.

Unknown is not failed.

Return to lesson
Idempotency, retries, and timeoutsThe order was saved, but the reply was lost. What makes the retry safe?Recall first, then reveal

The same authenticated caller, request key and matching request recover the saved order and result. A changed request using that key must be rejected.

Same request → same saved result.

Return to lesson
Idempotency, retries, and timeoutsSafe retriesRecall first, then reveal

Safe operation + bounded budget + backoff/jitter + reconciliation.

Retry with a reason and a limit.

Return to lesson

Final revision

Summary and interview notes

Retry the same logical operation only when its effect can be recovered safely and the remaining budget justifies another attempt. A durable operation identity prevents duplicate local mutations; also define how external results are checked, how obsolete workers are prevented from publishing, and how long saved results remain available for retries.

Remember these points

  • Bind the key to the caller, operation, target and normalized request fields. Reuse it when retrying the same request.
  • Save the request claim, business change and result in one atomic transaction.
  • A timeout leaves the result unknown. After ownership changes, the store must reject results from the former worker.
  • Deduplication retention and provider key lifetime bound safe retry; an expired record can make an old request look new.
  • Three retrying layers with three total attempts each can create 27 downstream calls.

Interview tips

  • Draw the crash after commit but before reply, then add two concurrent retries.
  • Show how a duplicate returns an already committed result before re-running create-time validation.
  • Count automatic SDK/proxy retries and include connection, queue and backoff time in the deadline.

Important qualifications

  • Saved outcomes still require current resource authorization.
  • Cancellation and circuit breakers reduce future work; they do not reverse an external effect already committed.

Technical references

Practice marks stay in this browser.