System-design interview · Extended interviews
Design a webhook delivery platform
Design how a service sends saved event notifications to customer URLs, signs each request, retries failures within limits, and helps receivers avoid repeating the business action.
You will learn to
- Separate event identity, subscription, delivery, attempt, and receiver processing.
- Trace a lost HTTP response without losing the event or claiming exactly-once effects.
- Design fair endpoint scheduling, signing, replay protection, and operational recovery.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Message queues, event logs, delivery guarantees, and backpressure · Design a distributed message log · Databases, data models, and ACID transactions · Production readiness: SLI, SLO, observability, and recovery
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Problem and scope
A webhook platform delivers signed HTTP events to customer-controlled endpoints. The sender records what happened and each delivery attempt; the receiver must save the incoming event and separately perform its business action, such as notifying the customer. Define a successful acknowledgement precisely, and bound attempts to unreachable endpoints. Event E402 reports shipment version 7 for O901. NorthHarbor may save it while the sender sees only a timeout because the 202 reply was lost.
Smallest working design
Begin with one database and a worker: commit the shipment and a pending notification record together, then let the worker POST it. The pending record survives a restart. Calling the customer inside the shipment transaction would hold locks across an uncontrolled network request and still leave uncertainty if the response vanished.
The sender’s pending notification record is an outbox: durable work saved with the shipment change. The receiver uses a separate inbox to remember events it has accepted. These records solve opposite sides of the handoff—avoiding a forgotten send and recognizing a repeated receipt—and cannot be replaced by one shared HTTP success flag.
Clarify the delivery contract
Candidate: “Does success mean NorthHarbor has received the event or completed its shipment workflow?” Interviewer: “Durably accepted the event; their workflow can run later.” Candidate: “Can their endpoint remain offline?” Interviewer: “Yes; retry for a bounded period and show terminal failures.” These answers prevent an impossible promise of guaranteed delivery to an indefinitely unreachable receiver.
Failure case to prove
Test a reply lost after the receiver saves the event. The sender keeps the event and attempt history and reports the outcome as uncertain. On retry, the receiver recognizes the same event despite the new HTTP attempt.
02Functional requirements
- Manage endpoints. Register, update or delete endpoint URLs; choose event-type subscriptions; rotate signing secrets. Authorize every action within its tenant.
- Deliver signed events. Attempt delivery at least once during the declared retention window. A receiver's 2xx means durable acceptance; eventual business processing is a separate outcome.
- Retry within bounds. Retain event bodies for seven days in this exercise. Retry eligible failures with exponential backoff and jitter; keep terminal failures inspectable.
- Inspect delivery state. Show pending, in-flight, retrying, accepted, exhausted, paused and canceled. Distinguish a known 400 response from an unknown timeout and show the next retry time.
- Redrive retained events. Permit an authorized operator to request delivery again, called a redrive, while preserving the original event ID and recording the reason.
- Offer scoped ordering. Default to best-effort order with event identity and object version for reconciliation. Optional per-object ordering sends one object's events in sequence; if an earlier event cannot be accepted, later events for that object wait. This delay is head-of-line blocking.
- Isolate subscribers. One broken endpoint must not stop unrelated customers.
Configuration changes and scope
A delivery records its target endpoint/configuration version. Changing a URL must not silently reinterpret a historical attempt. Here queued deliveries retain their planned version unless an explicit, audited redrive selects a new configuration.
There is no global total order across tenants and objects, nor a sender-only promise of exactly-once business processing. If the product needs to track the receiver's eventual workflow success, add a separate status protocol.
Stripe documents concrete duplicate/out-of-order webhook deliveries; that is a provider example, not a universal guarantee. Stripe webhooks.
03Non-functional requirements
- Workload assumption. Ten million business events/day, with three matching endpoint subscriptions/event.
- Latency and availability. For healthy endpoints, start the first attempt within five seconds of committing the event for at least 95% of deliveries. Target 99.95% scheduler availability, measured as the share of time it can claim and dispatch due work under the documented load.
- Durable acknowledgement. A successful business mutation and its outbox event survive one database-node failure together. The service must not report durable acceptance merely because it placed work in an in-memory queue; a process failure would erase it.
- Retention and retry budget. Shared payloads survive seven days; delivery/attempt audit metadata follows an explicit retention policy. Bound both attempt count and event age: the first exhausted limit produces diagnostics and exhausted state.
- Fair resource limits. Enforce per-endpoint concurrency, per-tenant fair capacity and global socket limits. Bound connection/response time and captured response bytes. Use an illustrative five-second HTTP timeout and delayed retries.
- Security per attempt. Recheck signing and destination validity for every attempt, so queued work cannot bypass a revoked endpoint or changed destination.
Correctness invariants
| Boundary | Required guarantee |
|---|---|
| Planning | One logical delivery per tenant, event, endpoint and configuration version |
| Retries and redrive | Stable event identity across all attempts |
| Worker ownership | A stale worker cannot overwrite newer delivery state |
| Tenant isolation | No cross-tenant event or secret exposure |
| Receiver processing | Local business effects require an inbox/idempotency contract; the sender cannot guarantee exactly once |
A permanently failing receiver cannot be guaranteed to accept an event. Eventual attempted delivery and eventual successful processing are different promises. A slow large customer must not consume every worker.
04Capacity estimates
Assume 10M business events/day, three matching endpoints/event, and one KB/event body.
| Quantity | Calculation | Consequence |
|---|---|---|
| First attempts | 10M × 3 / 86,400 ≈ 347/s | Average before retry traffic |
| Tenfold peak | 347 × 10 ≈ 3,470/s | Scheduler and connection capacity |
| In-flight at 0.5s average | 3,470/s × 0.5s ≈ 1,735 requests | Bound sockets/timeouts |
| Seven-day shared payload | 10M × 1 KB × 7 = 70 GB | Store payload once, reference per delivery |
| One-hour outage backlog | 3,470/s × 3,600s ≈ 12.5M attempts due | Stagger recovery, do not stampede |
Separate payload bytes from delivery work
Persist delivery metadata separately from payload bytes. Per-endpoint concurrency and tenant budgets are as important as total worker count; a large slow subscriber should not consume the entire connection pool.
Transport concurrency
If each logical delivery averages 1.2 attempts, average transport load is about 417 attempts/s and the tenfold peak is about 4,167/s. At a five-second timeout, an all-slow peak could occupy over 20,000 sockets, much more than the 1,735 normal in-flight estimate. Concurrency limits therefore enforce a capacity budget independently of arrival rate.
Storage and retry backlog
At 30M logical deliveries/day and an illustrative 200 bytes of base metadata, seven days is 42 GB before indexes/attempt history; payload sharing avoids storing the same 1 KB event three times. If a day's deliveries average two retained 150-byte attempt records, attempt metadata adds about 9 GB/day. Measure actual row/index overhead and avoid saving arbitrary response bodies indefinitely.
Recovery time
Draining a 12.5M-attempt backlog at an extra 1,000 attempts/s takes about 3.5 hours, assuming receivers can accept that load. Recovery cannot occur instantly by adding workers if endpoint limits are the bottleneck. Use per-tenant fairness and staggered due times so a recovering subscriber does not starve fresh healthy traffic.
05APIs and contracts
The envelope carries three levels of identity: E402 is the shipment event, D22 is its planned delivery to one endpoint, and A2 is one HTTP attempt. Retries change the attempt and signing timestamp while preserving the event being delivered. Keeping those identities separate lets inspection explain repeated network calls without inventing repeated shipments.
POST /webhook-endpoints
{url:"https://customer.example/events",eventTypes:["shipment.shipped"]}
→ {endpointId:EP9,configVersion:3,status:"active"}
Outbound body:
{id:E402,type:"shipment.shipped",schemaVersion:2,
tenant:NorthHarbor,objectId:O901,objectVersion:7,data:{...}}
Headers: delivery-id D22; attempt A2; timestamp; key-id; signature
Signature and retry identity
Document precisely which bytes/fields are signed and how timestamps/key IDs are encoded. Sign the exact body sent on the wire; a receiver verifies the raw bytes before parsing or normalizing JSON. The signing timestamp changes for a fresh attempt while E402 and D22 remain stable. Attempt IDs identify transport observations, not new shipment events.
Inspection and redrive
Inspection lists delivery state, next retry, attempt timestamps, response class and safe correlation IDs with an opaque cursor. It does not expose another tenant's payload or secret. POST /deliveries/D22/redrives requires authorization and reason, returns a redrive record and preserves event identity. Once the retained payload has been deleted, the service cannot reproduce the original event; respond with a clear unavailable-history result.
Configuration conflicts and tenant authorization
Endpoint changes use expected configuration versions to prevent lost edits. Registration validates URL syntax and ownership process as appropriate, but every connection still validates the resolved destination. A successful registration does not establish that future DNS answers are safe.
06Data model and access patterns
| Record/API | Example |
|---|---|
| Business event | E402,order=O901,type=shipment.shipped,objectVersion=7 |
| Subscription | EP9,tenant=NorthHarbor,eventTypes=[shipment.shipped],configVersion=3 |
| Logical delivery | D22,tenant=NorthHarbor,event=E402,endpointId=EP9,configVersion=3,state=pending |
| Transport attempt | A1,delivery=D22,attempt=1,startedAt,status=timeout |
| Receiver receipt | Inbox(E402,receivedAt,processingState=queued) |
| Registration | POST /webhook-endpoints {url:...,eventTypes:[...]} |
Planning uniqueness
Use a unique (tenantId,eventId,endpointId,configVersion) constraint when planning deliveries. A configuration version is local to its endpoint: E402 sent to EP9/version3 and EP10/version3 requires two distinct deliveries. Omitting endpointId would collapse them into one row and lose a destination. Replanning E402 for the same endpoint and version returns the existing delivery; retrying D22 retains that delivery identity. A new attempt must not look like a new shipment event. Schema versions keep old payloads interpretable. Keep event payloads immutable. Specify whether each event contains a snapshot of the object when the event occurred, or only identifies an object that the receiver must fetch in its current state.
Authoritative versus derived state
The business database stores O901 and outbox E402 atomically. An outbox is a database record of work to publish, committed in the same transaction as the business change so a crash cannot preserve one without the other. A payload store or event table keeps immutable event bytes with schema version/checksum. Delivery rows are partitioned by tenant/endpoint for fair scheduling and indexed by (state,nextAttemptAt). Each delivery includes retry count, current lease token, endpoint configuration reference and terminal reason. Each attempt records what the sender observed, within size limits. It does not replace or modify the saved business event.
Secrets and history
Secret references point to a controlled secret store, with key IDs and activation/retirement periods; plaintext keys never enter ordinary telemetry. A due scheduler can enqueue delivery IDs into a ready queue, but the delivery database remains authoritative when a queue item is repeated or lost. Periodic due scans repair missed publication.
Receiver-owned inbox
The receiver's inbox is separate infrastructure under NorthHarbor's control. Key it by trusted sender/tenant/event identity, store a payload fingerprint and processing state, and reject a conflicting payload with the same identity. If their inbox retention is shorter than our permitted redrive horizon, their one-effect guarantee ends early; this must be agreed rather than assumed.
07Basic working design
Commit, plan and attempt
Begin with one database and worker.
- Commit the event. The shipment transaction changes O901 to shipped and inserts E402 in an outbox.
- Plan the delivery. A planner reads E402, matches EP9 and inserts D22 under a unique tenant/event/endpoint/configuration-version constraint.
- Lease, send and record. A periodic due scan finds D22, records attempt A1 and a lease, sends a signed POST and records the observed outcome. The lease gives one worker temporary ownership; its token lets the database reject updates from a worker whose ownership has expired or been replaced.
Interpret the result
On a 202 response, D22 becomes accepted. On a timeout, the sender cannot tell whether NorthHarbor received the bytes; it records unknown transport outcome and schedules another attempt according to policy. It does not roll back shipment O901 or create a new shipment event. Once the business transaction commits, a worker can retry the saved delivery without repeating that transaction.
Keep remote I/O outside locks
This baseline can serve a small product safely if it has bounded timeouts, destination checks and persistent state. It already handles a process restart because due work remains in the database. The worker must not hold a transaction or row lock while making the remote HTTP call. It records a lease in a short transaction, releases locks, then uses a guarded result update afterward.
Receiver acknowledgement rule
The shipment transaction ends before calling the customer; pending delivery survives process failure.
Read each connection in order
- sync1. Commit O901 + E402Shipment application → Shipment / outbox / delivery DB
- async2. Lease pending D22Shipment / outbox / delivery DB → Due delivery worker
- sync3. Signed POST E402Due delivery worker → Customer HTTPS endpoint
- sync4. 2xx or uncertain timeoutCustomer HTTPS endpoint → Due delivery worker
- sync5. Save result if lease still currentDue delivery worker → Shipment / outbox / delivery DB
08Find the baseline flaws
| Failure test | What breaks and what must follow |
|---|---|
| Throughput and noisy endpoint | A single blocking worker at 0.5 seconds/request handles only two attempts/s, far below the 347/s average. A five-second slow endpoint reduces it to 0.2/s and blocks unrelated customers. Increasing threads without per-endpoint limits makes a large failing subscriber occupy the entire pool. This is the first scaling problem. |
| Lost acceptance response | The critical correctness test is: NorthHarbor inserts E402 into its inbox and commits, then the 202 response is lost. The sender times out and retries. If the receiver performs its shipment action before checking a unique inbox record, it can send two customer notifications or double-update a balance. If the sender refuses to retry, an alternative history in which the first request never arrived loses the event. There is no transport-only choice that distinguishes those two histories. |
| Expired worker writes late | A second failure is stale worker state. A1 times out locally, its lease expires and A2 succeeds. The old A1 worker resumes and blindly writes retrying, undoing accepted. Guarding updates with the current lease token and terminal-state rules prevents that local corruption, but cannot stop the remote receiver from seeing both POSTs. The design therefore needs both sender state fencing and receiver business deduplication. |
09Improve the design, step by step
Retry eligibility answers whether another attempt could help; retry timing answers when to make it. Exponential backoff increases the delay after repeated failures, and jitter varies that delay across deliveries so many workers do not retry together. Both remain bounded by the attempt and retention budgets already promised to the customer.
| Response/outcome | Example policy | Reason |
|---|---|---|
| 2xx | Mark accepted | Receiver contract acknowledged receipt |
| Timeout/network/5xx | Bounded backoff with jitter | Potentially transient or uncertain |
| 429 | Honor bounded retry guidance plus backoff | Receiver is controlling load |
| Permanent endpoint/configuration error | Pause or fail with diagnostics | Blind retries may never help |
Document which response statuses trigger retries. Index deliveries by their next attempt time. A worker claims a lease and records the attempt before releasing it. Per-object first-in, first-out (FIFO) delivery can delay later events behind a poison delivery: an event that repeatedly fails; parallel delivery improves throughput but requires object versions or receiver reconciliation. A timestamp alone is not a reliable total order. A receiver can fetch current object state when older events arrive late.
1. Bounded parallel workers with endpoint limits
- Trigger: the two-attempt/s baseline.
- Mechanism: A due scheduler leases many independent deliveries but enforces endpoint/tenant/global concurrency. Healthy tenants gain throughput without letting one slow endpoint own all sockets.
- Benefit, cost and alternative: Costs include fairness state and distributed limits; independent worker-local limits can multiply the cap. Keep one worker when volume is tiny and isolation unnecessary.
2. Durable ready queues and retry timing
- Trigger: database due scans or outage backlogs dominate.
- Mechanism and tradeoff: Publish delivery IDs into partitioned queues and use a durable due-time index for delayed retries. This reduces polling load and smooths recovery. A queue message may repeat or never arrive. Workers check the current delivery row before sending, and periodic scans enqueue pending deliveries that were missed. The queue tells workers which deliveries to inspect; the database still records pending work and can reconstruct a missing queue entry.
3. Payload sharing and immutable configuration references
- Trigger: three endpoint copies/event and audit ambiguity after URL edits.
- Mechanism: Store E402 once, reference it from D22, and pin endpoint/schema versions. This saves bytes and makes attempts explainable.
- Benefit, cost and alternative: It adds payload-store reads and retention coordination; garbage collection cannot delete bytes still needed by permitted retries. Inline payload rows remain simpler at small scale.
4. Receiver inbox and explicit ordering options
- Trigger: duplicate transport and out-of-order updates.
- Mechanism: Document atomic inbox acceptance and idempotent processing; optionally serialize deliveries per object when needed.
- Benefit, cost and alternative: This improves business correctness at the cost of receiver storage or head-of-line blocking. An object-version/current-state fetch model is preferable when strict sequence is unnecessary and recovery speed matters.
10Detailed architecture
Sender authority and worker fleet
The sender's business transaction owns O901 and E402. A planner creates logical delivery rows from subscriptions and immutable payloads. A due scheduler feeds ready work to a bounded worker fleet, with delivery state and lease tokens checked in the authoritative database. A signing component accesses secret references, and an egress policy layer validates the actual destination before HTTP is sent.
Receiver transaction boundary
The receiver is outside our trust and transaction boundary. It verifies sender signature, validates schema and commits E402 to its own inbox before returning 2xx. Its processing worker then applies the business effect under its own idempotency/transaction rules. There is no arrow claiming one transaction spans our delivery row and their shipment database.
Inspection and exhaustion
Metrics and inspection use sanitized attempt history. When retries are exhausted, the delivery database retains the failed record, reason and permitted redrive actions. A dead-letter queue, if used, must preserve a link to that inspectable record. Per-endpoint pause/deletion policy is checked before a new attempt, even when an old queue message exists.
Trust boundaries in the diagram
The final diagram deliberately shows secret storage and egress enforcement because this service makes requests to customer-controlled URLs. A generic worker-to-internet arrow would hide a material trust boundary. Delivery acceptance and receiver processing are also separate boxes so an interviewer can point to exactly which 202 acknowledgment is being discussed.
A local sender transaction cannot include the receiver. The receiver acknowledgment follows its own durable inbox commit.
Read each connection in order
- sync1. Commit mutation + eventBusiness event producer → Business database / outbox
- async2. Read committed eventBusiness database / outbox → Subscription delivery planner
- syncStore/reuse immutable event bytesSubscription delivery planner → Immutable event payloads
- sync3. Unique event/endpoint/versionSubscription delivery planner → Delivery / attempt authority
- async4. Schedule due delivery IDDelivery / attempt authority → Due / ready work queue
- async5. Deliver ready delivery IDDue / ready work queue → Fair leased delivery workers
- syncClaim token / record attemptFair leased delivery workers → Delivery / attempt authority
- syncRead pinned E402 bodyFair leased delivery workers → Immutable event payloads
- syncResolve active signing keyFair leased delivery workers → Signing secret store
- sync6. Signed bounded requestFair leased delivery workers → Destination checks and egress
- sync7. Safe HTTPS POSTDestination checks and egress → Customer webhook endpoint
- sync8. Unique inbox + durable jobCustomer webhook endpoint → Receiver inbox / job store
- sync9. 2xx after inbox commitCustomer webhook endpoint → Fair leased delivery workers
- sync10. Save result if lease still currentFair leased delivery workers → Delivery / attempt authority
- async11. Apply local idempotent effectReceiver inbox / job store → Receiver business worker
- syncInspect / audited redriveTenant inspection / redrive API → Delivery / attempt authority
11Write path and acknowledgement
The shipment commit and dispatch intent survive together. Attempts preserve the logical event identity while using separate transport-attempt identities.
Numbered delivery trace
- Commit business change and outbox. The shipment transaction changes O901 to shipped/version 7 and writes outbox E402.
- Plan one delivery. The planner creates D22 for EP9/version 3 exactly once; its due time is now.
- Claim and send. A leased worker records attempt A1, signs the raw E402 payload plus a fresh timestamp, and sends it.
- Receiver accepts; reply is lost. NorthHarbor verifies the signature, inserts inbox E402 durably, and responds 202. The response is lost.
- Retry the same identity. The sender records an uncertain timeout and schedules A2 with the same event/delivery identity.
- Deduplicate at the receiver. NorthHarbor recognizes inbox E402, does not enqueue another shipment effect, and returns 202 again.
- Record acceptance. D22 becomes accepted. NorthHarbor’s worker independently completes its idempotent business update.
The sender did not learn whether A1 arrived. Stable identity plus receiver persistence makes that ambiguity recoverable.
Claim and revalidate the attempt
Before sending A2, the worker claims D22 with a new lease token and persists the attempt start. It reads the pinned event/configuration, checks that delivery is still allowed, creates a fresh timestamp/signature and opens a bounded connection through destination validation. No database transaction remains open during this request.
Fence the result update
After the response, a guarded update requires delivery.leaseToken == myToken and the expected in-flight state. A2's valid 202 can set accepted and append attempt details atomically. If the lease changed, the worker appends only an appropriately associated observation or returns stale-attempt; it does not overwrite the current delivery state. All captured data is bounded and sanitized.
Two distinct deduplication boundaries
The planner's unique delivery key and the receiver's inbox key protect different boundaries. The first prevents duplicate planned subscriptions; the second prevents duplicate downstream effects. Neither means that only one TCP connection or POST ever occurred. That distinction is the central interview answer.
12Read and delivery path
Status distinguishes known acceptance, scheduled retry and unknown outcome. Redrive respects receiver deduplication and retention.
Numbered inspection and redrive flow
- Authenticate the dashboard query. NorthHarbor opens its delivery dashboard. Authentication establishes the tenant; the query uses tenant plus delivery ID and a bounded attempt-history cursor.
- Show known and unknown outcomes. The API returns D22's current state, the pinned endpoint version, last observed response class, next attempt and event retention deadline. An A1 timeout is labeled uncertain, not definitively rejected.
- Authorize redrive. A user requests redrive with a reason. The service verifies payload retention, endpoint eligibility and authorization, creates an audited redrive request, and schedules the original event identity under the chosen configuration policy.
- Apply ordinary delivery guards. The worker follows the same signature, destination, lease and fair-capacity path as automatic delivery. Manual actions do not bypass egress checks or tenant limits.
- Report acceptance without inventing a new effect. The receiver may recognize E402 as already processed and immediately return 2xx. The dashboard then records accepted again without claiming a new business shipment occurred.
Due-time scheduling
Retry scheduling reads a due-time index ordered by deadline, not a loop scanning every historical attempt. A leased item is skipped until its lease expires or completes. Backoff with jitter prevents a one-hour outage from turning into millions of synchronized POSTs. Inspection can use replicas for older history, but status after a redrive should reflect its committed request/version or clearly indicate lag.
13Correctness deep dive
Atomic receiver acceptance
NorthHarbor first verifies the signature and schema. It then performs one short transaction:
transaction receive(trustedSender, trustedTenant, eventId, rawBody):
scope = (trustedSender, trustedTenant, eventId)
inserted = insert inbox(scope, hash(rawBody), state=QUEUED)
if absent under unique(scope)
lock inbox[scope]
require inbox[scope].payloadHash == hash(rawBody)
if inserted: insert unique processing_job(scope)
commit
return HTTP 202
transaction process(scope):
lock inbox[scope]
if state == DONE: commit; return
apply local business mutation for the same trusted tenant
set inbox[scope].state = DONE
commit
Race outcomes
Receiver commits first: A1 creates inbox E402 and its job, commits and loses the response. A2 races with processing, finds the same inbox identity and returns 202. The unique insert plus transaction ensures only one durable job; the processing transaction ensures a crash cannot commit the business mutation without the DONE state when both share that database.
Receiver crashes before commit: no inbox/job exists, so A2 inserts them and proceeds. Processing crashes after commit: the next job sees DONE and makes no second effect. Two workers processing the same event serialize on the inbox row.
External effects need another boundary
Fence sender state separately
Sender fencing is separate: UPDATE deliveries SET state=accepted WHERE id=D22 AND leaseToken=L2 AND state=in_flight. A stale L1 cannot undo L2's accepted result. Retain inbox identity at least through the sender's allowed replay/redrive horizon, or explicitly accept that older manual replays require business-level duplicate detection.
Composite identity and payload conflicts
The inbox, processing job, retry lookup and business update all identify the event by verified sender, tenant and event ID together. Event ID alone is insufficient. Those values come from the endpoint's verified signing-credential mapping, not an arbitrary unsigned tenant header. Concurrent receipt transactions either observe the committed existing row or retry a uniqueness/serialization conflict; neither creates a second job. A matching event ID with different bytes is rejected rather than silently treated as a duplicate.
Two HTTP attempts produce one durable inbox identity and one local business effect under the stated transaction contract.
Read each connection in order
- syncA1 POST E402 / D22Sender worker → Receiver endpoint
- syncInsert E402 + job; commitReceiver endpoint → Receiver inbox DB
- blocked202 response lostReceiver endpoint → Sender worker
- syncA2 POST same E402 / D22Sender worker → Receiver endpoint
- syncUnique lookup finds E402Receiver endpoint → Receiver inbox DB
- return202 acceptedReceiver endpoint → Sender worker
- syncLock E402; apply local effect + DONEReceiver processor → Receiver inbox DB
- returnCommit effect and DONEReceiver inbox DB → Receiver processor
- syncRecord D22 accepted under leaseSender worker → Sender worker
14Failure and recovery
| Failure or condition | Surviving state, response and recovery |
|---|---|
| Worker dies after sending | If a worker dies after sending but before recording success, its lease expires and the same delivery retries. A fenced lease token prevents stale workers overwriting newer attempt state; it does not stop a remote endpoint from seeing duplicates. Keep receiver inbox/business mutations idempotent. Stripe API idempotency keys illustrate a provider-scoped retry contract, but they are distinct from deduplicating incoming webhook event IDs. Stripe idempotent requests. |
| Manual redrive or endpoint deletion | Manual redrive preserves the original event identity and records an operator/redrive reason. A receiver whose dedupe retention is shorter than the sender’s replay window can repeat effects; align those contracts or require explicit replay-aware processing. After endpoint deletion, cancel future delivery according to an auditable policy. |
| Lost outbox/queue publication | If the business service commits O901 but crashes before outbox publication, the relay resumes from the durable outbox. If the planner creates D22 but loses its queue publish, the due scan recovers it. If the worker sends and crashes before recording an outcome, its lease eventually expires and the same event retries. Each boundary has surviving state rather than a generic “retry everything” instruction. |
| Subscriber outage | During a subscriber outage, use endpoint-specific backoff and a circuit/pause policy while continuing other tenants. A 429 may carry retry guidance; bound and validate it so a malformed value does not retain work forever. Status-code classification is documented because not every 4xx is safely permanent for every integration. |
| DNS, key or payload changes | If DNS changes from a public address to an internal destination between attempts, egress validation rejects the new attempt without contacting it. If signing keys rotate while old attempts remain pending, use the configured overlap/key-ID contract rather than silently signing with an unknown key. If payload retention expires, mark exhausted/history-unavailable explicitly; don't reconstruct an old snapshot from today's object state and call it the same event. |
15Operations, security, and cost
Signature verification and replay policy
A valid signature establishes that the body came from a party holding the signing key and was not altered. The receiver must still validate business fields and authorize the requested operation. Sign the exact transmitted bytes with a documented timestamp/key identifier, verify against the raw request body, compare securely, and reject excessive timestamp age under a replay policy. Rotate secrets with a bounded overlap; store secret references rather than plaintext in logs. Fresh signatures on retries are compatible with stable event IDs.
Endpoint and tenant security
Endpoint URLs create a server-side request-forgery surface. Restrict schemes, validate resolved destinations at connection time, block internal/metadata addresses, and disable redirects or validate every hop. Registration-time DNS checks alone do not handle later DNS changes. Authenticate endpoint changes and do not expose another tenant’s events through delivery inspection.
Latency and backlog metrics
Observe first-attempt success, eventual acceptance, oldest pending age, attempts per logical delivery, per-endpoint sockets, queue delay and lease expiry. Tie healthy first-attempt latency to the five-second objective and separately report retry backlog; a high eventual success rate can hide hours of delay. Capture only bounded sanitized response snippets because receivers may return secrets or customer data.
Socket and storage costs
At the illustrative slow peak, 20K sockets plus TLS buffers can dominate worker memory even though payloads are only 1 KB. Per-endpoint concurrency one limits a five-second failing endpoint to roughly 0.2 active attempts/s before backoff; extra workers cannot responsibly drain that endpoint faster without changing policy. Shared payloads save about 140 GB over seven days compared with three independent 70 GB copies, before replicas, under the given assumptions.
Rollout and failure drills
Roll out envelope/schema changes additively with versioned contracts and test receivers. Drill sender death after POST, receiver death after inbox commit, expired lease late writes, DNS rebinding, duplicate redrive and one tenant's hour-long outage. A successful drill proves one local business effect under the inbox assumptions while acknowledging repeated HTTP transport.
Bind DNS validation to the actual socket
16Decision ledger and limitations
Expose outcomes without overclaiming
Expose pending, retrying, accepted, exhausted, and paused states with attempt timestamps, sanitized response codes, and correlation IDs. Limit response-body capture: endpoints may return sensitive content. Measure first-attempt success, eventual acceptance, oldest pending age, per-endpoint backlog, retry amplification, signature failures, worker lease expiry, and receiver latency percentiles.
Failure drill
Run a failure drill: persist E402, kill the sender after POST, restore it, then kill NorthHarbor after inbox commit. Show one final business effect despite repeated transport. Next pause EP9 for an hour and prove other tenants retain capacity. “We retry” is incomplete unless retention, identity, fairness, and visibility are demonstrated together.
Decision table
| Decision | Benefit | Cost/limit | Change trigger |
|---|---|---|---|
| At-least-once attempts plus inbox | Recovers uncertain delivery | Receiver must persist dedupe state | No generic transport-only exactly-once alternative |
| Per-endpoint fair concurrency | Isolates slow subscribers | Backlog may drain slowly | Receiver explicitly accepts higher parallelism |
| Shared immutable event payload | Efficient, reproducible retries | Retention coordination and payload reads | Tiny volume favors simpler inline rows |
| Best-effort order plus versions | Parallel throughput | Receiver reconciliation | Workflow truly requires serialized per-object delivery |
| Seven-day retry horizon | Bounded storage/work | Long outages become terminal | Product funds a longer documented replay window |
Ordering and replay-retention limits
Strict per-object FIFO makes a poison event block later updates. Skipping it restores throughput but changes the ordering contract; inspect and decide, rather than silently doing both. Snapshot events preserve historical facts, while thin notifications followed by a current-object fetch simplify convergence but may omit intermediate states. Choose based on what the subscriber needs to do, not only payload size.
The sender controls its saved attempts and reported outcomes. The receiver must uphold its promise to save work before returning 202; an external provider must supply any duplicate-safe behavior its effects require. Integration guidance and contract tests are therefore part of the system design, not optional documentation after the worker code is done.
17Interview closing
Rehearse the architecture and contract
“I commit the business change and its outbox event together, plan one delivery per tenant, event, endpoint and configuration version, and let leased workers attempt delivery outside the business transaction. Retries preserve event and delivery identities but create new attempt identities and signatures. Fair endpoint budgets and durable due scheduling keep one outage from consuming the fleet. Before accepting a worker’s update, the database checks its lease token; a late worker cannot overwrite newer delivery state.
Defend the critical boundary
“The hard network case is a receiver commit with a lost 202. We must retry, so the receiver verifies the signature and atomically stores a unique inbox event plus processing work before acknowledging. Its local effect and DONE state commit together, or external effects use their own idempotency protocol. The costs are duplicate transport, retained inbox/delivery state and bounded terminal failures. My next tests are lost replies at both commit boundaries and an hour-long noisy endpoint outage.”
Answer the follow-up
If the interviewer demands ordering, scope it per object or subscription and explain poison-event blocking. If they demand exactly-once processing across a third-party payment call, explain the missing shared transaction and design an explicit operation identity/reconciliation boundary rather than promising that a message queue setting solves it.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What does the receiver’s 202 response mean in this design?
Reveal a model answer
It means the event is durably accepted into its inbox, not that all shipment side effects have finished. That lets the endpoint respond quickly without losing work after a crash. If the sending service needs proof of completed processing, I would define a separate status or callback contract.
Interviewer follow-up
Why not acknowledge before storing the inbox row?
Reveal the follow-up answer
A crash in that gap loses the event while the sender believes delivery succeeded. Persisting the receipt first makes either retry or internal recovery possible.
What the answer must demonstrate: State what the receiver has saved before returning 202.
A1 arrived, but the sender never saw its response. What changes in A2?
Reveal a model answer
The event E402 and logical delivery D22 stay the same. The attempt number, send timestamp, and signature are new. NorthHarbor deduplicates the stable event identity and returns acceptance again without repeating the business effect.
Interviewer follow-up
Can sender-side dedupe avoid the second network call?
Reveal the follow-up answer
No. The sender cannot infer the lost response’s outcome. It needs a retry or a supported receipt-query protocol; receiver-side idempotency resolves repeated arrival.
What the answer must demonstrate: Uncertain transport requires cooperation at the receiver.
Shipment version 8 arrives before version 7. Should the receiver roll back its state?
Reveal a model answer
For a full versioned snapshot, the receiver atomically installs only a newer object version, so version 7 cannot replace version 8. For dependent deltas, it detects the missing sequence and replays or fetches an authoritative complete state instead of silently discarding earlier work. Sender ordering through 202 controls receipt order; the receiver must separately order processing if required.
Interviewer follow-up
Can event-created timestamps replace versions?
Reveal the follow-up answer
Not reliably. Different events can share timestamps and clocks/transport can reorder them. Use a documented ordering token or an explicit state-reconciliation rule.
What the answer must demonstrate: State which field orders events and how the receiver handles an older snapshot or a missing delta.
One large customer’s endpoint stalls for thirty seconds per request. What protects others?
Reveal a model answer
Per-endpoint concurrency limits, connection timeouts, and tenant scheduling budgets prevent that customer from occupying every worker/socket. A durable due-time queue retains its backlog, and jittered retries avoid a synchronized recovery flood when it returns.
Interviewer follow-up
Why not launch more workers without limits?
Reveal the follow-up answer
They can amplify load against the receiver and exhaust our sockets or spend. Capacity expansion does not replace fairness and a bound on outstanding work.
What the answer must demonstrate: Reason about in-flight requests as well as request rate.
Why verify the raw body instead of parsed and reserialized JSON?
Reveal a model answer
A signature authenticates specific bytes. Reserialization can alter spacing, field order, or number formatting even when the parsed object appears equivalent, causing verification failure. I verify the original body under the documented signature/timestamp scheme before trusting its contents.
Interviewer follow-up
Does a valid signature prevent repeated effects?
Reveal the follow-up answer
No. A legitimate old event can be replayed within a permitted window. Timestamp policy limits replay exposure, while event identity and an inbox protect the business mutation.
What the answer must demonstrate: Authenticity and idempotency solve different problems.
An operator redrives E402 three months later. Can it safely reuse the same event ID?
Reveal a model answer
The chosen service retains event payloads for seven days, so a three-month redrive is rejected as unavailable history. If a separate archival contract retains the original payload longer, preserve its event identity and audit the redrive, but align receiver deduplication or business reconciliation with that extended horizon. Reconstructing today’s object is not replaying the original event.
Interviewer follow-up
What if the endpoint changed tenants meanwhile?
Reveal the follow-up answer
Endpoint ownership/subscription versions and authorization must be checked. A recycled URL or mutable tenant mapping must not receive historical events belonging to another customer.
What the answer must demonstrate: Replay safety includes lifetime and ownership, not only a UUID.
A2 succeeds, then the old A1 worker reports timeout. How do you prevent accepted becoming retrying?
Reveal a model answer
The delivery row stores the current lease token and state. A result update must carry that token, so A1’s old token cannot replace A2’s accepted result. Keep A1’s late observation in attempt history without changing the delivery outcome.
Interviewer follow-up
Does this stop A1 from reaching the customer twice?
Reveal the follow-up answer
No. Remote transport can still duplicate. The receiver inbox and business idempotency contract handle that separate boundary.
What the answer must demonstrate: Distinguish sender-state fencing from receiver deduplication.
The receiver inbox transaction is safe, but processing sends a payment. Is the payment exactly once?
Reveal a model answer
Not from the inbox transaction alone. The payment service is external, so persist a stable outgoing operation identity/outbox and use its idempotency/status contract. Reconcile uncertain outcomes before creating another financial operation.
Interviewer follow-up
What if the provider has a shorter idempotency retention window?
Reveal the follow-up answer
Our retained attempt state must prevent blind retries outside that window; use status reconciliation or an explicit recovery policy. Retention is part of the end-to-end guarantee.
What the answer must demonstrate: Do not extend a local transaction across a network call.
Blank-page exercise · 45 minutes
Build the answer yourself
Deliver the sending service’s E402 shipment event to EP9. Lose A1’s response, crash both sender and receiver at different points, and then redrive an old event after secret rotation.
- Define event, delivery, attempt, and inbox identities.
- Calculate fanout, in-flight requests, and outage backlog.
- Trace durable receipt before acknowledgement.
- Choose retry, ordering, and fairness contracts.
- Verify raw bytes and restrict outbound destinations.
- Reconcile retry retention with historical redrive.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a webhook delivery platformWhat does a successful webhook response prove?Recall first, then reveal
The receiver accepted the request under its documented contract; it does not necessarily prove the downstream business job finished.
Accepted is not completed.
Return to lessonDesign a webhook delivery platformWhat remains stable across retries?Recall first, then reveal
The event and logical delivery identity; attempt number, request timestamp, and signature can change.
Same event, new attempt.
Return to lessonDesign a webhook delivery platformWhy should receivers persist before acknowledging?Recall first, then reveal
Otherwise a crash after 2xx but before durable enqueue loses an event the sender considers delivered.
Save receipt, then say received.
Return to lessonFinal revision
Summary and interview notes
The sender saves delivery work and retries within limits; the receiver saves receipt and protects its local business transaction from duplicates. Stable event identity and lease-token checks let both recover from lost replies, even though HTTP requests may repeat.
Remember these points
- Commit business state and its outbox together; retain pending deliveries independently of the ready queue.
- Event and delivery identity survive retries, while attempt identity, timestamp and signature change.
- Use the same trusted sender/tenant/event key for receipt, job and processing; local effect and DONE commit together.
- Waiting for each 202 can order receiver acceptance; the receiver must separately order processing if later business actions depend on earlier ones.
- Endpoint fairness, retry age and payload retention bound cost and recovery promises.
Interview tips
- Draw both indistinguishable timeout histories: the POST never arrived, or receipt committed and the response vanished.
- Separate snapshot version handling from delta ordering, and distinguish sender fencing from receiver deduplication.
- Trace the exact resolved address used for the connection, not just a registration-time URL check.
Important qualifications
- The five-second timeout and seven-day payload window are exercise choices, not Stripe guarantees.
- Historical redrive beyond retained payloads is unavailable unless a separate archive contract exists.
- External receiver effects still need their own idempotency and reconciliation protocol.
Technical references
- Stripe webhook documentationProvider-specific examples of signature verification, duplicate deliveries, event ordering, and retry behavior.
- Stripe idempotent requestsDistinguishes outbound API retry keys from an application’s webhook inbox and business deduplication policy.
- OWASP SSRF prevention guidancePrimary security guidance on destination validation, unsafe address ranges and redirect handling for outbound requests to user-controlled URLs.
Practice marks stay in this browser.