System designby Learnastra

Concept lesson · Foundations

Message queues, event logs, delivery guarantees, and backpressure

By Anup Rai

Start here

Definition

A message queue buffers work for asynchronous consumers. An event log retains an ordered history for consumers to read or replay. Backpressure controls admission or processing concurrency when downstream capacity cannot keep up with incoming work.

Why it matters: Slow background processing should not hold every foreground request open. Buffering absorbs short bursts, while durable handoff and duplicate-safe processing make accepted work recoverable.

The visual modelQueue backlog growth and drain time

A queue absorbs a burst but does not create service capacity. Admission limits and bounded retries prevent a growing backlog from becoming an outage.

Queue backlog growth and drain timeA queue absorbs a burst but does not create service capacity. Admission limits and bounded retries prevent a growing backlog from becoming an outage. At 600 jobs per second arriving and 400 per second completed, backlog grows by 200 each second. After sixty seconds it grows by 12,000. To drain an existing backlog, completion capacity must exceed arrival rate. Acknowledge a delivered job only after its durable effect. A crash may cause redelivery, so effects need stable identities.A queue buys time; workers provide throughput600 jobs/sarrive400 jobs/sfinishBacklog grows by 200 jobs/ssecondsqueued jobs060+12,000growth = arrivals - completionsAt 200 arrivals/s and 400 completions/s, the 12,000-job backlog drains in 60 s.
Read the diagram step by step
  1. At 600 jobs per second arriving and 400 per second completed, backlog grows by 200 each second. After sixty seconds it grows by 12,000.
  2. To drain an existing backlog, completion capacity must exceed arrival rate.
  3. Acknowledge a delivered job only after its durable effect. A crash may cause redelivery, so effects need stable identities.

Worked example

Workers process 400 jobs/s while 600 jobs/s arrive for 60 seconds: the backlog grows by 12,000 jobs. When arrivals fall to 200/s, the spare 200 jobs/s drains it in about 60 seconds.

Key takeaways

  • Accepted work and completed work are different user-visible states.
  • At-least-once delivery requires a safe repeated effect; an outbox prevents lost handoff, not duplicates.
  • A queue stores excess work but cannot fix sustained overload without more capacity or less admission.

You will learn to

  • Separate accepting a job from completing its business effect.
  • Trace the database-to-queue gap and a crash after the effect but before acknowledgment.
  • Compute backlog growth and recovery while bounding retries and resource use.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Databases, data models, and ACID transactions · Replication and durability

Workload and timing examples are interview assumptions.

01What is a message queue, and what does async mean?

A message queue holds work until a consumer can process it. The sender is the producer; a broker is the service that stores and delivers the messages. Asynchronous means the request can finish accepting work before that work finishes executing. Backpressure is the control that slows, defers, or rejects incoming work when processing capacity is insufficient.

An event records something that happened, such as photo-created. A command asks for an action, such as create-thumbnail. A queue commonly distributes commands among workers; a retained log allows independent consumers to replay events. These uses can share infrastructure, but their completion and replay contracts differ.

Synchronous processing makes request latency include the entire downstream task and occupies request-serving capacity throughout it. For example, accepting photo P501, rendering a thumbnail, and publishing its ready state can have very different latency distributions. A burst of slow rendering jobs can exhaust request workers even when upload storage is fast.

A queue stores work so another process can handle it later. We create job J501 and return a status saying the uploaded file and a record that it needs thumbnail processing have been durably saved. This is not the same as “thumbnail ready.” The client receives a photo ID and can check states such as pending, processing, ready, or failed.

The queue decouples when work arrives from when it executes. It can absorb a bounded burst, but it cannot make sustained excess demand disappear. The first design decision is therefore a user-visible contract: acceptance is quick and recoverable; completion is asynchronous and has a separate objective. All job rates and timestamps below are illustrative assumptions.

02Work queue versus event log versus publish/subscribe

A work queue distributes tasks among workers. J501 should be handled by an eligible thumbnail worker, with retry if that worker fails. A visibility lease can temporarily hide the task from other workers, but expiration can lead to another delivery while the first worker still runs.

Concept in focusWho receives the work?

Arrows show delivery; upward arrows under the retained log mark independent reader positions.

Who receives the work?Arrows show delivery; upward arrows under the retained log mark independent reader positions. Trace a job to one worker, a log to two reader positions, and an event to two subscriptions. Workers compete for J1, J2 and J3 in the work-queue example. Readers A and B can be at different positions in E1 through E4. Subscribers A and B each receive E1; their delivery guarantees depend on the implementation.Work queue: deliver different jobs to competing workersJ1J2J3 waitsWorker A: J1Worker B: J2Event log: each reader has its own positionE1E2E3E4Reader AReader BPublish-subscribe: subscriptions each receive a copyE1Subscriber A: E1Subscriber B: E1

Remember: Queue: divide jobs. Log: retain history. Pub/sub: distribute copies.

Read the diagram
  1. Trace a job to one worker, a log to two reader positions, and an event to two subscriptions.
  2. Workers compete for J1, J2 and J3 in the work-queue example.
  3. Readers A and B can be at different positions in E1 through E4.
  4. Subscribers A and B each receive E1; their delivery guarantees depend on the implementation.
Try from memoryWhich picture lets two readers replay the same retained history at different speeds?

The event log with independent reader positions. A competing-worker queue instead divides jobs among workers.

A retained event log stores an ordered sequence that consumers can replay from a position. Separate consumer groups can independently process the same photo events: one builds thumbnails, another computes usage statistics. Publish/subscribe describes sending events to multiple subscribers; durability, retention, and replay depend on the actual system.

Mechanism P501 use Question to answer
Work queue Assign thumbnail job J501 When is it eligible for retry?
Retained log Replay photo-created events How long are events retained?
Publish/subscribe Notify independent consumers Does each subscriber receive durable work?

For each system, specify whether messages can repeat, which messages stay ordered, how long history remains available, and what counts as completed work. The names “queue” and “publish/subscribe” do not promise global order or exactly-once effects.

For this thumbnail service, a coherent starting implementation is a PostgreSQL photo/outbox transaction, a retrying relay, an SQS standard work queue, and workers that commit result metadata back to PostgreSQL. The database stores the job's logical state; the queue schedules attempts. A retained partitioned log such as Kafka is useful instead when several consumers need independent replay of photo events. Its order is per partition, so choosing photo ID as a partition key does not provide one global order across every photo.

03Transactional outbox: avoid a lost database-to-broker handoff

Suppose the upload service first commits P501 and then sends J501 to the broker. It crashes between those actions. The photo exists, but no worker learns that processing is required. Reversing the order creates another gap: a job may exist for a photo record that never committed.

Concept in focusPut the business change and event in one commit

The shaded boundary is the local database transaction. Publication happens outside it.

Put the business change and event in one commitThe shaded boundary is the local database transaction. Publication happens outside it. Locate the atomic boundary and the later, retryable publication path. Order O17 and outbox event E17 commit together. A relay publishes E17 to the consumer. Uncertain publication may repeat; the consumer must deduplicate E17 with its effect.ONE DATABASE TRANSACTIONOrder O17: paidOutbox: E17RelayConsumerpublish E17Commit both rows, or neither.Relay may publish twice. Consumer deduplicates E17 with its effect.

Remember: Commit the order and outbox together; expect relay retries.

Read the diagram
  1. Locate the atomic boundary and the later, retryable publication path.
  2. Order O17 and outbox event E17 commit together.
  3. A relay publishes E17 to the consumer.
  4. Uncertain publication may repeat; the consumer must deduplicate E17 with its effect.
Try from memoryCan the relay publish E17 twice even though the database committed once?

Yes. It can lose a publication confirmation and retry. The consumer needs a durable duplicate guard coupled to its effect.

A transactional outbox puts the photo record and a row describing the job to be published, including its ID and payload, in one database transaction. They commit together. A relay reads committed outbox rows and publishes jobs. If the broker is unavailable, the intention remains durable for a later retry. This protects the handoff without pretending the database and broker share one local transaction. Outbox reference.

The relay marks an intention published only after the broker confirms the required durable acceptance. A timeout is an unknown outcome, so it retries with the same event ID. Marking the row first would recreate the lost-handoff gap. Monitor oldest unpublished-outbox age separately from broker queue age: work can be stuck before it ever reaches the queue.

The relay needs a way to discover newly committed outbox rows. It can repeatedly query the table, or follow the database’s committed change stream. Change data capture provides the second option; a saved checkpoint records publication progress so a replacement connector can resume.

Change data capture (CDC) exports committed database changes to downstream systems, often by decoding the transaction log. A connector takes a consistent snapshot, continues from its matching log position, and checkpoints progress. If a crash occurs after publishing but before checkpointing, the connector can publish a change again; consumers still need replay-safe writes.

CDC can publish an outbox table without application polling. Capturing every table update instead is a different contract: low-level row changes do not necessarily represent a complete business event. Define transaction boundaries, keys, deletion records and schema evolution. In PostgreSQL, a stalled logical replication slot can retain WAL and exhaust storage, so monitor retained bytes as well as connector lag. CDC does not make an external effect atomic with the source transaction.

Interview check: Why retain both an outbox and CDC? The outbox defines the business event within the source transaction; CDC is one transport for publishing it.

Worked example diagramPhoto P501 and job intent J501 commit together. The relay can publish duplicates. A worker writes immutable attempt output, then a unique job receipt, current-version check, and authoritative reference share one transaction before queue acknowledgment.
Message queues, event logs, delivery guarantees, and backpressure: architecture diagram1. Accept upload P501 to 2. Photo + outbox transaction: accept original and processing intention; 2. Photo + outbox transaction to 3. Relay: read committed outbox; 3. Relay to 4. Durable queue J501: publish; duplicates possible; 4. Durable queue J501 to 5. Thumbnail worker: lease/deliver J501; 5. Thumbnail worker to 6. Result reference + unique receipt: atomic receipt, version check, and reference; 6. Result reference + unique receipt to 4. Durable queue J501: acknowledge completed work; 6. Result reference + unique receipt to 7. Ready status exposed: status becomes ready1 → 2: accept original and processing intention2 → 3: read committed outbox3 → 4: publish; duplicates possible4 → 5: lease/deliver J5015 → 6: atomic receipt, version check, and reference6 → 4: acknowledge completed work6 → 7: status becomes ready01Accept upload P50102Photo + outboxtransaction03Relay04Durable queue J50105Thumbnail worker06Result reference +unique receipt07Ready status exposed
  1. 1 → 2accept original and processing intentionAccept upload P501 → Photo + outbox transaction
  2. 2 → 3read committed outboxPhoto + outbox transaction → Relay
  3. 3 → 4publish; duplicates possibleRelay → Durable queue J501
  4. 4 → 5lease/deliver J501Durable queue J501 → Thumbnail worker
  5. 5 → 6atomic receipt, version check, and referenceThumbnail worker → Result reference + unique receipt
  6. 6 → 4acknowledge completed workResult reference + unique receipt → Durable queue J501
  7. 6 → 7status becomes readyResult reference + unique receipt → Ready status exposed

04Consumer acknowledgments, duplicate delivery, and idempotent effects

A consumer acknowledgment tells the broker that a delivery has been handled and can be marked complete under the queue’s contract. The worker must choose that moment carefully: acknowledging before its result is recoverable can lose work after a crash, while acknowledging later allows duplicates that must be safe to handle.

The worker receives J501, whose immutable intent identifies photo P501, its version, and the thumbnail recipe. It creates an immutable output object for this attempt, then commits the authoritative output reference and a receipt for J501 in one database transaction. The receipt is a deduplication record: evidence that this logical operation has a recorded outcome. Only after that commit does the worker acknowledge the queue message.

Concept in focusAcknowledgement must follow the durable effect

Commit the result and its duplicate-detection record together. Acknowledging the message before committing the result can lose work after a crash.

Acknowledgement must follow the durable effectCommit the result and its duplicate-detection record together. Acknowledging the message before committing the result can lose work after a crash. Broker to Consumer: Deliver event E. Consumer to Effect store: Commit E's effect with a durable duplicate guard. Consumer to Broker: Acknowledgement is lost after the effect commits. Broker to Consumer: Redeliver E after uncertain acknowledgement. Consumer to Effect store: Recognize the existing effect and avoid applying it again.BrokerConsumerEffect storeDeliver event E.Commit E's effect with a durable duplicate guard.Acknowledgement is lost after the effect commits.Redeliver E after uncertain acknowledgement.Recognize the existing effect and avoid applying it again.

Remember: Effect first, acknowledgement second, replay safely.

Read the diagram
  1. Broker to Consumer: Deliver event E.
  2. Consumer to Effect store: Commit E's effect with a durable duplicate guard.
  3. Consumer to Broker: Acknowledgement is lost after the effect commits.
  4. Broker to Consumer: Redeliver E after uncertain acknowledgement.
  5. Consumer to Effect store: Recognize the existing effect and avoid applying it again.
Time Event Recovery implication
12:00:00.000 P501 and outbox J501 commit Relay can recover the job
12:00:00.005 Worker receives J501 Job is not yet complete
12:00:00.080 Thumbnail exists; ready record and receipt commit Effect is recoverable
12:00:00.090 Worker crashes before queue acknowledgment Redelivery is expected
Later New worker sees J501 receipt Return existing outcome and acknowledge

The object-store write and the database transaction still commit separately. The stable logical identity is photo/version/recipe; immutable attempt-specific object keys avoid two concurrent renders overwriting the same bytes. A protected database transaction chooses the authoritative reference. A database receipt alone cannot make an unrelated external API call atomic.

Cleanup must coordinate with the transaction that makes the output available to readers. An unreferenced object may belong to an active render that has not committed yet. Keep an attempt record protecting it; cleanup first marks an expired attempt abandoned under the same transactional state that publication checks. An abandoned attempt cannot subsequently publish. Delete only abandoned, unreferenced attempt objects, so a scan that observed no reference cannot race with a later valid commit.

05At-most-once, at-least-once, and ordering guarantees

At-most-once handling can avoid repeated attempts by discarding or acknowledging before the effect, but a crash can lose work. At-least-once delivery permits repeats so incomplete or uncertain work can be attempted again. For example, SQS standard delivery explicitly requires duplicate-aware applications.

For J501 we choose at-least-once delivery with an idempotent effect: repeated processing converges on the same recorded thumbnail result. Retain deduplication evidence for the supported retry/replay horizon. Reusing the same job ID for different photo contents must fail or follow a defined versioning rule.

Ordering also needs a scope. If P501 version 3 replaces version 2, a late version-2 job must not overwrite the version-3 ready record. A conditional version check protects that update. Partitioning events by photo can help order their handling, but retries and parallel execution still require a precise rule for applying results.

Delivery/effect promise What happens after ambiguity Remaining application responsibility
At most once No redelivery after the chosen discard/ack boundary Accept possible lost work
At least once Delivery may repeat Deduplicate the effect and define retry/retention limits
Exactly once within a transaction boundary Result and consumed position commit together Verify the boundary includes every promised effect

In a retained log, a consumer’s offset is its recorded position in a partition. Atomically committing that position with new output records prevents one of those facts from advancing without the other. This is useful when consuming one event produces another, but the transaction still has a defined storage boundary.

06Backpressure: calculate queue growth and drain time

Concept in focusThe queue grows, then drains

Time runs left to right; height is the number of waiting jobs. Rates are constant within each 30-second interval.

The queue grows, then drainsTime runs left to right; height is the number of waiting jobs. Rates are constant within each 30-second interval. Read the rise and fall of a 600-job backlog. For 30 seconds, 120 arrivals/s minus 100 completions/s adds 600 jobs. Then 80 arrivals/s leaves 20 completions/s for the backlog. It drains in another 30 seconds.Backlog = arrivals minus completions, accumulated over time600 jobs00 s30 s60 s+20 jobs/s-20 jobs/s120 in; 100 out80 in; 100 outAfter arrivals fall, 600 / (100 - 80) = 30 seconds to drain.

Remember: Drain time uses spare capacity, not the full service rate.

Read the diagram
  1. Read the rise and fall of a 600-job backlog.
  2. For 30 seconds, 120 arrivals/s minus 100 completions/s adds 600 jobs.
  3. Then 80 arrivals/s leaves 20 completions/s for the backlog. It drains in another 30 seconds.
Try from memoryWhy does draining take 30 seconds instead of 6?

New arrivals still consume 80 of the 100 completions/s. Only 20/s is available for the 600 waiting jobs: 600 / 20 = 30 seconds.

If arrivals remain at 400/second, the backlog does not drain. If they stay above 400, it grows. A queue is a buffer, not additional processing capacity.

Backpressure means controlling incoming or concurrent work when downstream capacity cannot keep up. Bound queue size or accepted delay, limit per-tenant demand, and reject or defer new uploads before exhausting durable storage. Retry temporary failures with bounded attempts and randomized backoff. Send permanently invalid images to an explicit failed/dead-letter workflow with diagnostics; retrying malformed bytes forever wastes capacity. Track oldest-job age as well as queue length, because it measures completion delay.

A dead-letter destination retains work that exhausted its retry policy or needs intervention, together with enough failure information to inspect it. An operator or recovery process may correct the cause and replay it deliberately. Moving a job there records an unresolved or failed outcome; it does not complete the thumbnail.

Limit work already taken by workers: fetching millions of jobs only moves the backlog into their memory. Cap concurrent render tasks and extend visibility timeouts while valid long jobs run; duplicates remain possible. After an outage, limit replay speed so old retries leave capacity for new work. If job sizes differ greatly, limit estimated resource use as well as job count.

07Interview answer: define the exactly-once effect boundary

Interviewer: “Can this queue guarantee thumbnails are processed exactly once?”

Candidate: “I would first distinguish rendering the thumbnail from publishing the result that users can see. J501 can be delivered again after a worker commits the ready record but crashes before acknowledging. I would make output publication safe to repeat and record J501’s receipt with the database result, so the retry returns the established outcome.

“I would use an outbox to ensure every accepted photo P501 has a stored job recording the required processing. That relay can also repeat publication, so duplicate handling remains necessary. For overload, I would measure job age and bound admission. A 600-per-second burst cannot be solved indefinitely by workers that process 400 per second.”

This answer traces both handoffs: accepting work into the background system and committing the worker’s result. It defines observable pending, ready, and failed states. The idempotency-retries-and-timeouts and distributed-transactions-and-workflows chapters develop the related boundaries further.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What is a message queue? Define producer, consumer, asynchronous completion, and backpressure.

Reveal a model answer

A producer submits a message; a broker stores and delivers it; a consumer processes it. A message queue buffers that work so its execution can happen asynchronously, after the initial request has durably accepted it. Accepted and completed are different states. A photo upload can return a job ID while its thumbnail is still pending.

Backpressure limits incoming or concurrent work when downstream processing cannot keep up. If arrival is 600 jobs/s and workers can process 400 jobs/s, the backlog grows by 200 jobs/s. Bound admission or increase processing capacity instead of treating an unbounded queue as a solution. For reliable handoff, an outbox can commit the photo and job intention together; consumers still need duplicate-safe effects because delivery or acknowledgment can repeat.

What the answer must demonstrate: Acceptance and completion are different promises.

Foundation · Question 2

How does a work queue differ from a retained event log?

Reveal a model answer

“The work queue assigns J501 and manages retry eligibility. A retained log lets consumers replay an ordered sequence from their positions, possibly in independent groups. I would choose based on work assignment versus replay needs and inspect the real retention and delivery semantics.”

What the answer must demonstrate: Do not infer guarantees from product-category names.

Applied · Question 3

A database commits photo P501, then the process dies before publishing its job J501. How does a transactional outbox close that gap?

Reveal a model answer

“I store the photo and an outbox intention in one transaction. A relay can find committed unpublished intentions after recovery. This prevents P501 being accepted without a durable record that processing must happen, while avoiding a false claim of one transaction across database and broker.”

What the answer must demonstrate: Outbox solves a missing handoff, not every duplicate.

Applied · Question 4

The worker commits the result and crashes before acknowledging. Walk the retry.

Reveal a model answer

“The job is delivered again after the broker did not record completion. A worker reads the durable job receipt and returns the established outcome. If two duplicate workers race before either receipt exists, a unique job-ID constraint and one result transaction choose the winner; the loser rolls back and reads that winner’s outcome. A preflight receipt lookup alone would not prevent two concurrent effects.”

What the answer must demonstrate: Deduplication placement determines correctness.

Follow-up · Question 5

Does a consumer’s local deduplication receipt make an external API call exactly once?

Reveal a model answer

“No. A remote object write is outside the result database transaction. I give each render attempt an immutable object key and use the stable photo/version/recipe identity for the logical job. The database transaction chooses one reference and records its receipt; it never lets a duplicate overwrite the chosen bytes. Cleanup cannot delete an active attempt that may still publish. A different provider effect needs that provider’s idempotency contract or reconciliation of unknown outcomes.”

What the answer must demonstrate: Local atomicity does not automatically include a remote effect.

Applied · Question 6

A thumbnail job for photo version 2 finishes after version 3 is published. What should the result commit check?

Reveal a model answer

“The authoritative ready update checks which photo version it belongs to. A stale version-2 result may be retained or cleaned up, but it cannot replace the version-3 reference. Ordering events by photo can help, yet the conditional update protects against retry and completion reordering.”

What the answer must demonstrate: Stable identity must represent stable intent.

Applied · Question 7

For 60 seconds, arrivals are 600 jobs/s and workers complete 400/s. Arrivals then fall to 200/s. Calculate backlog growth and drain time.

Reveal a model answer

“Six hundred arrivals minus four hundred completions gives two hundred extra jobs per second. Over sixty seconds that is twelve thousand jobs. When arrivals fall to two hundred, the spare capacity is two hundred, so recovery takes about sixty seconds under stable-rate assumptions.”

What the answer must demonstrate: Use service rate minus arrival rate when estimating drain time.

Follow-up · Question 8

J501 contains an image that can never be decoded. Should it retry forever?

Reveal a model answer

“No. I would classify the permanent failure, stop after a bounded policy, store diagnostics, and expose a failed state or dead-letter workflow. Infinite retries consume capacity and can delay valid work. Any operator replay should be controlled and preserve job identity and version rules.”

What the answer must demonstrate: Bound both retry effort and user-visible delay.

Blank-page exercise · 18 minutes

Build the answer yourself

Trace job J501 through upload acceptance, outbox publication, thumbnail creation, and queue acknowledgment. Crash one component between every adjacent pair of steps.

  • Store a processing job for every accepted photo.
  • Commit the selected thumbnail reference and its unique job receipt in the same database transaction.
  • Explain what happens after a permanent processing failure.
  • Calculate backlog and drain time for the stated burst.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Message queues, event logs, delivery guarantees, and backpressureWhat does the outbox guarantee for P501?Recall first, then reveal

The photo record and intention to publish J501 commit together; a relay can retry publication after failure.

Save the photo record and pending job together.

Return to lesson
Message queues, event logs, delivery guarantees, and backpressureWhy can J501 arrive again after success?Recall first, then reveal

The worker may commit its result but fail before the broker records acknowledgment.

Result committed, receipt on the wire lost.

Return to lesson
Message queues, event logs, delivery guarantees, and backpressureHow long does a 12,000-job backlog take to drain at capacity 400 and arrivals 200 per second?Recall first, then reveal

About 60 seconds under stable-rate assumptions: 12,000 divided by 200 spare jobs/second.

Drain with spare capacity, not total capacity.

Return to lesson
Message queues, event logs, delivery guarantees, and backpressureWhy should a version-2 job not overwrite version 3?Recall first, then reveal

Delivery order and completion order can differ; condition the authoritative update on the intended photo version.

An old job cannot replace a newer photo version.

Return to lesson

Final revision

Summary and interview notes

A durable queue separates acceptance from execution and absorbs bounded bursts. Reliable completion depends on recovering the database-to-broker handoff, preventing a repeated job from changing the published result twice, and keeping admitted work within processing and storage capacity.

Remember these points

  • A transactional outbox commits the business record and publication intention together; the relay can still publish duplicates.
  • Commit the result and a uniquely constrained job receipt together before acknowledging delivery.
  • Immutable attempt outputs and one authoritative reference prevent duplicate renders from overwriting selected bytes.
  • Ordering is scoped, often per key or partition; a version check prevents old work replacing newer state.
  • Backlog drain uses spare capacity: 12,000 jobs / (400 − 200 jobs/s) = 60 seconds.

Interview tips

  • Crash the producer between database and broker, then crash the consumer between result commit and acknowledgment.
  • Make two workers check an absent receipt simultaneously; show the unique constraint or transaction that resolves the race.
  • Ask what an exactly-once claim includes: broker records, a database effect, or an external provider.

Important qualifications

  • Visibility timeout is a retry mechanism, not proof the previous worker stopped.
  • Object cleanup must exclude attempts still authorized to publish; checking for a missing reference once is insufficient.
  • Retention, retry limits, and dead-letter policy bound recovery; permanent invalid jobs need an explicit failed outcome.

Technical references

Practice marks stay in this browser.