System designby Learnastra

System-design interview · Core interviews

Design a ticket-booking service

By Anup Rai

Design all-or-nothing seat allocation, durable holds, fair admission and payment recovery; preserve inventory ownership when expiry, retries and late provider outcomes race.

You will learn to

  • Demonstrate all-or-nothing reservation of overlapping seat sets.
  • Separate a five-minute business hold from a short database lock.
  • Recover hold expiry, waiting order, and ambiguous payment outcomes.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Databases, data models, and ACID transactions · Quorums, consensus, leases, and fencing · Message queues, event logs, delivery guarantees, and backpressure · Caching: cache hits, misses, write policies and invalidation

Workload and timing examples are interview assumptions.

01Problem and scope

A ticket-booking service sells scarce seat inventory while allowing browsing, temporary holds and payment. Its core invariant is that one show-seat has at most one current allocation and a requested seat set commits entirely or not at all. For example, customer A requests seats 54–56 while customer B requests 56–57 for show S99. Both maps may show seat 56 as available, but the authoritative booking transaction must allow only one competing seat set to succeed.

Scope the product to exact seat selection with all-or-nothing requests of up to ten seats, a five-minute hold, and an explicit bounded payment-processing grace period. Include city/movie/cinema/show discovery, seat maps, payment confirmation, cancellation and durable waiting. Resale, cross-show carts and dynamic auction pricing are excluded. This is an interview design, not a claim about Ticketmaster’s internal implementation.

The payment-processing grace period is another server deadline: it bounds how long an eligible hold remains allocated while an initiated payment is being resolved. It does not make a payment timeout a decline or let repeated checkout requests extend the hold indefinitely.

02Functional requirements

  1. Search and browse: Search by city/postal code or coordinates/radius, keyword, date and showtime; sort/paginate and offer spelling help where useful. Browse movie → cinema → hall → show → seat map, including seat class and price.
  2. Hold an exact seat set: Success reserves every requested seat under one hold; conflict reserves none. A free-looking map seat remains advisory until allocation commits.
  3. Pay and recover status: Pay for a hold, observe pending/confirmed/expired, and retrieve the booking after a lost response. A payment timeout is not a definite decline.
  4. Cancel a hold: Cancel when product rules permit.
  5. Wait for admission: Join, inspect or leave a durable per-show waiting list while capacity is temporarily held; notify the browser when admission or booking status changes.

Guest ownership and waiting limits

Guest checkout uses a server-issued unguessable session credential, not a caller-supplied name. A returning guest recovers the same hold through that credential. Waiting ends on cancellation, a one-hour session limit or a provably impossible request. Zero free seats alone does not prove impossibility because existing holds can expire.

Strict FIFO admission protocol

Admission controls who may attempt to create a hold when too many customers are competing. An admission grant is temporary permission to make that attempt; it does not itself reserve seats. The hold transaction must validate and consume the grant while allocating the requested seat set.

First-in-first-out (FIFO) admission serves waiting customers in their persisted order. A show lane is the queue of requests competing for the same show inventory; this strict version allows only one customer in that lane to hold an unused admission grant at a time.

  1. Select the oldest eligible ticket: Persist waiting order and allow at most one unconsumed admission grant per contending show lane.
  2. Issue a short grant: The hold endpoint must validate that active grant.
  3. Consume and allocate atomically: The hold transaction consumes the grant with seat allocation. Later grant expiry cannot release a hold already created from it.
  4. Advance the lane: Admit the next ticket only after the prior grant is consumed, canceled or expires.
Admission choice Consequence
Strict FIFO A six-adjacent-seat request can block smaller requests when only scattered singles remain.
Configured bypass Improves utilization but changes the fairness promise.
Multiple simultaneous grants FIFO issuance does not ensure FIFO redemption or strict seat-allocation priority.

A wider admission window is an explicit higher-throughput fairness alternative. Explain this choice rather than promising both strict FIFO and maximum utilization.

03Non-functional requirements

  1. Read latency: Catalog/seat-map p95 below 200 ms; the advisory seat map may be stale for up to two seconds.
  2. Booking latency: Admitted hold creation p95 below 500 ms; status updates within two seconds of a committed transition. Requests past their deadline fail clearly rather than queue indefinitely.
  3. Availability: 99.95% monthly availability for admitted booking requests. Inventory safety takes priority over accepting writes during an authority partition.
  4. Durability: A successful hold survives one database-node failure. Use three replicas across failure domains and an explicitly configured durable commit policy; replica count alone does not establish a recovery point objective (RPO), the maximum acknowledged-data loss allowed after a failure.
  5. Regional recovery: Assume a tested 15-minute restoration objective and disclose any asynchronous cross-region recovery-point gap.
  6. Retention and payment security: Size booking/payment audit records for an assumed five years, subject to business policy. Keep payment credentials at the provider. Expired sessions and wait tickets need much shorter retention.

Inventory and payment invariants

Invariant Required behavior
Exclusive seat allocation (showId,seatId) has at most one current allocation.
Atomic seat set An order gets every requested seat or none.
Terminal hold state An expired/released hold cannot steal seats back.
Stable provider identity One logical payment attempt uses one provider identity.
Authoritative confirmation Neither a browser redirect nor a stale map approves a reservation.

These are interview assumptions to negotiate, not measured product facts. A latency/availability target must not weaken the seat-allocation contract.

04Capacity estimates

Workload assumptions and arithmetic

Assume 3 billion page views and 10 million ticket sales per month. At 30 days/month, these are about 1,157 views/s and 3.86 tickets/s, not orders/s. At two tickets/order, the average is only 1.93 orders/s. That modest average must not justify an unbounded checkout endpoint for a popular on-sale event.

Worked estimates

Quantity Calculation Consequence
Daily show-seat inventory 500 cities × 10 cinemas × 2,000 seats × 2 shows 20 million rows/day
Raw inventory storage 20M × 100 bytes × 1,825 days 3.65 TB before indexes and replicas
Illustrative burst 100,000 shoppers opening one show in 10 s 10,000 seat-map reads/s
Map payload 10,000/s × 50 KB 500 MB/s if every response reaches origin
Hold admission budget 500 successful/failed attempts/s × 20 ms mean service About ten concurrent active transactions before lock waiting

Capacity implications and limits

The last number is an initial benchmark target, not a claimed database capability. Contention on the same seat can dominate even when CPU is idle. Cache catalog and advisory maps; keep admission proportional to measured lock service capacity. At a five-minute hold lifetime, 500 admitted holds/s could create 150,000 active holds, although the available seats of one show impose a much smaller practical ceiling. Across many shows, index (state, deadline) for incremental expiry scans.

Store physical seat geometry and movie/hall metadata once. Repeating a hall layout inside every show-seat row inflates both cache and storage estimates. Price snapshots and the exact selected seats do belong with the order so later catalog edits cannot change an already-agreed charge.

The 500-attempt/s figure is an authority load-test budget, not the achieved rate of the strict FIFO waiting lane. One outstanding grant may be limited by client round-trip or its expiry interval, and a difficult head request can leave inventory idle. Measure that separately; obtaining higher throughput by issuing many concurrent grants changes the fairness contract.

05APIs and contracts

Interface contracts

API Successful result Important failure
GET /shows/S99/seats Layout, advisory states, asOf, map version A map never constitutes a hold
POST /shows/S99/holds 201 {holdId:H7, seats:[54,55,56], expiresAt, version:1} 409 seat_conflict, 429 admission_required
POST /holds/H7/checkout 202 {paymentAttempt:P3,state:processing,deadline} 409 expired or request-payload mismatch
GET /bookings/by-hold/H7 Pending, confirmed, expired or refund-recovery status Unauthorized ownership returns no private details
POST /shows/S99/wait-tickets Ticket W8 and current state Request exceeds show capacity or session limit

Validation and response semantics

The hold request contains Idempotency-Key: k7, seat IDs, an expected quoted price version and, when required, an admission grant. The server obtains session identity from authentication. Persist (showId, sessionId, key) with a canonical payload fingerprint and result in the same transaction as H7. Reusing k7 with another seat set is a conflict; retrying the same request returns H7 even after a response is lost. Specify retention long enough to cover the user retry window, and require explicit status recovery after that window.

The booking service must recover payment outcomes independently of the browser. A webhook is a provider-to-service notification delivered to a registered endpoint; it may arrive after the checkout request has timed out. The stored payment attempt connects that notification to the original hold.

The client supplies a new checkout request key, but the service creates and durably retains provider payment identity P3. A retried checkout reuses P3. Webhooks carry provider event IDs and require signature verification; duplicated or out-of-order callbacks must be reconciled against the durable attempt state. Search pagination uses an opaque cursor carrying filters and sort position, not an ever-increasing offset through a changing catalog.

06Data model and access patterns

ShowSeat tracks the allocation of one physical seat for one show; Hold groups the requested seat set, while PaymentAttempt tracks the separate external charge. RequestResult saves the response needed for retries. The outbox stores notification work in the same transaction as a booking change, so later delivery can recover from a crash.

Record Example key and fields Authority/query
ShowSeat (S99,56): state=held,holdId=H7,version=1 Booking authority; exact seat-set lookup
Hold (S99,H7): session=sessionA,state=held,deadline=12:05,version=1 Authority for lifecycle and ownership
HoldSeat (S99,H7,56) Immutable requested set, also 54 and 55
PaymentAttempt (S99,P3): hold=H7,providerId, outcome=unknown External outcome and reconciliation history
WaitTicket (S99,108): W8,requested=2,state=waiting,expiresAt=13:00 Durable FIFO order
AdmissionGrant (S99,G9): ticket=W8,consumed=false,deadline Required booking-path permission
RequestResult / Outbox (S99,session,key) / (S99,eventId) Replay response / recover notifications

Shard by ShowID, so seats, hold, admission and outbox for S99 share one relational transaction domain. Sharding by MovieID would put every showing of a blockbuster together without improving the seat invariant. City, Cinema, Hall, Seat, Movie and Show form the catalog relations; their materialized search index and seat-map cache are derived stores.

Create indexes on active holds by deadline, bookings by owner/show, waiting tickets by show/sequence and outbox entries by dispatch state. To validate H7, lock its hold row, load its immutable seat set, then lock seats in ascending ID order. Use the same order for every hold lifecycle operation. New hold creation locks requested seats in ascending order. Database constraints prevent duplicate request IDs; transaction logic enforces that all seats move together. An index cannot by itself encode every multi-row lifecycle rule.

For expiry, use the durable ordered deadline index to find due holds and an exact hold-ID lookup for cancellation/removal. A linked list ordered only by creation time is valid for expiry scanning only when creation order also preserves deadline order. Payment grace periods can change that order, so deadline indexing is the safer baseline. After worker restart, rebuild ready batches from stored active holds and waiting tickets rather than trusting a replicated in-memory list as the only recovery record.

For creation, a database unique constraint claims the request identity inside the same transaction as all seat effects. A conflicting duplicate rolls back any tentative work and reads the saved payload/result in a fresh transaction if required by the isolation mode. Checking for a missing request row before locking seats is not enough: both concurrent attempts can see absence, and the loser could otherwise report conflict after the winner already created its hold.

07Basic working design

One booking process and SQL authority

The first version has a browser, one booking process, a relational database and an external payment provider. Catalog reads, hold creation, expiry scans and payment reconciliation run in this process. This is enough to prove business behavior before introducing independent queues or caches.

Atomic all-seat hold creation

Customer A submits k7 for seats 54–56. The service authenticates the customer session, begins a transaction, claims the unique request-result key, and locks all three ShowSeat rows. A concurrent duplicate waits on that key and returns the committed matching result before attempting seat allocation; it does not report a seat conflict against its own successful first attempt. It verifies every seat is free and the quote is still valid. It inserts H7 and its HoldSeat rows, changes all three allocations, inserts the replay result, and commits. Only then does it return 201 H7. A crash before commit leaves no hold; a crash after commit but before response leaves H7 recoverable by k7.

Conflicts and payment intent

Customer B's request for 56–57 cannot partially succeed. If customer A owns the conflicting seat when customer B obtains the locks, the competing transaction rolls back. The customer can request alternatives or wait. The service does not lock seats while calling the provider: it first commits a bounded processing state and P3, then makes the network call.

Deployable baseline and limits

This baseline can run as a complete service: a periodic loop expires holds and reconciles payments from saved pending work, including after a restart. Its limitations are throughput, process availability and operational isolation; its basic inventory invariant is already valid.

architecture · baselineBaseline: prove one complete hold

One transaction changes the entire seat set. Payment remains an external boundary.

Baseline: prove one complete holdOne transaction changes the entire seat set. Payment remains an external boundary. client to api: 1. Hold k7 / seats 54–56; api to db: 2. Commit all-seat H7 transaction; db to api: 3. Durable H7 result; api to client: 4. Return hold and deadline; api to pay: 5. Pay only after P3 is durable1. Hold k7 / seats 54–562. Commit all-seat H7transaction3. Durable H7 result4. Return hold and deadline5. Pay only after P3 is durableACTORCustomer browserSERVICEBooking applicationSTORERelational bookingdatabaseEXTERNALPayment providersync
Read each connection in order
  1. sync1. Hold k7 / seats 54–56Customer browser → Booking application
  2. sync2. Commit all-seat H7 transactionBooking application → Relational booking database
  3. sync3. Durable H7 resultRelational booking database → Booking application
  4. sync4. Return hold and deadlineBooking application → Customer browser
  5. sync5. Pay only after P3 is durableBooking application → Payment provider

08Find the baseline flaws

Bottleneck / counterexample Evidence and design consequence
On-sale read and hold stampede First replay the on-sale workload: 10,000 map reads/s and a stampede of hold attempts all reach one process and one database. At 50 KB/map the process may try to emit 500 MB/s before TLS and query overhead. A popular seat causes waiting transactions that occupy connections; adding more web threads merely grows the waiting room inside the database. The performance bottleneck is partly repeated browsing work and partly serialized inventory contention. They need different changes.
Late payment after hold expiry Then inject a correctness failure into a tempting implementation. At 12:04:58 checkout calls the provider without first changing H7's durable state. At 12:05:00 expiry frees seats 54–56. Customer B then buys 56. At 12:05:01 the provider reports success for customer A. Code that blindly marks H7 confirmed oversells. Adding a cache or more replicas does not repair that interleaving.
Required durable payment state The correct baseline already needs a transaction that records processing, its deadline and P3 before the provider call; final confirmation must re-check both hold state and allocation ownership under locks. If expiry wins a later race, a successful external charge must create refund work, the compensating action for a charge without a booking. Test this by pausing the provider response, running expiry and then releasing the response. Record actual row values after each commit. This test defines the state machine that the larger design must preserve.

09Improve the design, step by step

  1. Change 1 — move advisory reads off the booking database. The trigger is the 500 MB/s map burst. Cache catalog and short-lived map snapshots behind the read service, while serving static geometry through an edge cache. At a 99% map hit rate, 10,000 requests/s become roughly 100 origin reads/s. The costs are cache memory, invalidation traffic and stale displays. The new failure is a convincing-looking stale map, so every hold still validates authority. Direct database reads remain preferable for a small deployment where freshness and simplicity outweigh cache savings.

  2. Change 2 — enforce durable admission. The trigger is lock queues exceeding the 500 ms hold objective. The wait coordinator grants a limited number of short-lived tokens in persisted sequence order; the hold transaction consumes the token. This bounds work reaching scarce seats and prevents newcomers bypassing waiting users. The cost is extra waiting latency, durable coordinator state and idle capacity when grants expire unused. The new risk is a restarted coordinator issuing stale grants: the database validates the active coordinator’s ownership version, called its epoch, when creating them. A simple reject-and-retry limit is cheaper when FIFO fairness is not a product promise.

  3. Change 3 — partition by show and replicate each authority. The trigger is aggregate inventory exceeding one database's measured write or storage capacity. A routing directory maps S99 to shard 12; all of its seat sets remain local transactions. Different shows gain parallel capacity; one hot show does not. Costs include rebalancing, replica lag and shard-aware operations. Stale routing can send requests to the old owner, so the handoff must prevent that shard from writing before the new owner accepts requests. A larger single database is preferable until this complexity buys measured headroom.

  4. Change 4 — extract durable background work. The trigger is payment timeouts and notification retries competing with interactive holds. Commit outbox records beside booking changes; dispatchers feed wait admission, map invalidation and browser notifications. Expiry and payment reconcilers use indexed durable work. This improves isolation and recovery, at the cost of duplicate events, lag and extra workers. Consumers deduplicate event IDs and fetch authoritative status. An in-process worker is still suitable while it has the same persistent work contract and adequate isolation.

10Detailed architecture

Read and booking request paths

The browser enters through an authentication/admission edge. Browsing goes to the catalog/map service and derived cache; booking writes go through the show router to the booking authority. The router consults a show-to-shard directory, which is configuration, not inventory. The final diagram separates these paths because a cached “available” response and a committed hold have different guarantees.

Show-local source of truth

Inside a show shard, the booking service and relational primary own ShowSeat, Hold, PaymentAttempt, waiting/admission state and outbox records. Replicas implement the promised durable commit and recovery policy. The application acknowledges success only after that policy completes. Other regions may serve catalog reads, but a single active authority accepts seat writes for S99. Promoting another owner requires fencing the previous writer through the storage failover mechanism; a DNS change alone is insufficient.

Expiry, reconciliation and notifications

Expiry/reconciliation workers and outbox dispatchers sit outside the synchronous response path. Each worker checks the saved hold or payment state in a transaction before changing it. The payment provider is an external transaction boundary. Its request and our local booking update cannot be wrapped in one database transaction. Browser notifications report changes; they never create allocations. The wait coordinator creates grants in the same authority that consumes them. If it cannot safely issue a grant, customers stay waiting while existing valid holds remain readable.

Hot-show transaction bottleneck

This architecture has an intentionally visible bottleneck: the transactional owner of one very popular show. Admission makes that limit survivable; replication and sharding do not abolish it.

architecture · finalFinal: isolate browsing, authority and recovery

Seat-map caches cannot grant inventory. Every hold, grant and lifecycle transition reaches the same show authority.

Final: isolate browsing, authority and recoverySeat-map caches cannot grant inventory. Every hold, grant and lifecycle transition reaches the same show authority. client to edge: 1. Authenticated browse / hold; edge to read: 2a. Browse S99; read to cache: Read advisory snapshot; read to route: Refresh map on cache miss; edge to route: 2b. Hold k7 with admission grant; route to dir: 3. Resolve S99 owner; route to book: 4. Route to authoritative shard; book to db: 5. Atomic seats + H7 + replay; db to rep: Replicate durable commit; book to pay: 6. Submit stable payment P3; pay to book: 7. Verified payment outcome; wait to db: Guard epoch; create active FIFO grant; reconcile to db: Guarded expiry / unknown outcomes; reconcile to pay: Reconcile P3 / refund identity; db to dispatch: 8. Read committed outbox; dispatch to cache: Invalidate snapshot version; dispatch to wait: Capacity / queue event; dispatch to client: Notify; client fetches status1. Authenticated browse / hold2a. Browse S99Read advisory snapshotRefresh map on cache miss2b. Hold k7 with admissiongrant3. Resolve S99 owner4. Route to authoritative shard5. Atomic seats + H7 + replayReplicate durable commit6. Submit stable payment P37. Verified payment outcomeGuard epoch; create activeFIFO grantGuarded expiry / unknownoutcomesReconcile P3 / refund identity8. Read committed outboxInvalidate snapshot versionCapacity / queue eventNotify; client fetches statusACTORCustomer browsersG1SERVICEIdentity / admissionedgeG1SERVICECatalog and mapserviceG2CACHEAdvisory map /catalog cacheG2SERVICEShow routerG2STOREShow -> sharddirectoryG2SERVICEBooking authorityG3STOREShow-shard primaryHolds / seats / outboxG3STOREDurable shardreplicasG3WORKERWait / grantcoordinatorG3WORKERExpiry / paymentreconcilerG3WORKEROutbox / statusdispatcherG3EXTERNALPayment providerG4syncreplicationasyncG1 Client / request boundaryG2 Derived browsing and routingG3 Show authority and recoveryG4 External payment boundary
Read each connection in order
  1. sync1. Authenticated browse / holdCustomer browsers → Identity / admission edge
  2. sync2a. Browse S99Identity / admission edge → Catalog and map service
  3. syncRead advisory snapshotCatalog and map service → Advisory map / catalog cache
  4. syncRefresh map on cache missCatalog and map service → Show router
  5. sync2b. Hold k7 with admission grantIdentity / admission edge → Show router
  6. sync3. Resolve S99 ownerShow router → Show → shard directory
  7. sync4. Route to authoritative shardShow router → Booking authority
  8. sync5. Atomic seats + H7 + replayBooking authority → Show-shard primary Holds / seats / outbox
  9. replicationReplicate durable commitShow-shard primary Holds / seats / outbox → Durable shard replicas
  10. sync6. Submit stable payment P3Booking authority → Payment provider
  11. sync7. Verified payment outcomePayment provider → Booking authority
  12. syncGuard epoch; create active FIFO grantWait / grant coordinator → Show-shard primary Holds / seats / outbox
  13. syncGuarded expiry / unknown outcomesExpiry / payment reconciler → Show-shard primary Holds / seats / outbox
  14. syncReconcile P3 / refund identityExpiry / payment reconciler → Payment provider
  15. async8. Read committed outboxShow-shard primary Holds / seats / outbox → Outbox / status dispatcher
  16. asyncInvalidate snapshot versionOutbox / status dispatcher → Advisory map / catalog cache
  17. asyncCapacity / queue eventOutbox / status dispatcher → Wait / grant coordinator
  18. asyncNotify; client fetches statusOutbox / status dispatcher → Customer browsers

11Write path and acknowledgement

A hold and a payment attempt are separate durable records. The example defines show S99, customer session sessionA, request k7, seats 54–56, hold H7, admission grant G9 and provider attempt P3 before tracing their commits.

  1. Customer A sends k7, quote version Q4, G9 and seats 54–56. The edge authenticates sessionA and routes S99 to shard 12; request limits stop repeated speculative holds.
  2. The transaction first claims the unique (S99,sessionA,k7) request row, serializing concurrent duplicates before grant consumption or seat locks. An existing identical request returns H7; a changed payload conflicts. It verifies and consumes G9 when waiting is active, locks seats in order and validates every allocation and quoted price.
  3. It writes H7 version 1, all seat allocations, the replay result and outbox event E71. After the durability policy succeeds, return the server deadline. A lost response is recovered with the same k7.
  4. Checkout locks H7 and its seats and first returns any matching saved checkout attempt. For a new checkout, if H7 is held and unexpired according to fresh authority time after lock waits, change it to processing version 2 with deadline 12:06, create P3 and commit. The provider call uses P3's stable idempotency identity and the stored amount/currency.
  5. A verified success follows the common order: H7, its seats in ascending order, then P3. If H7 is still valid processing and owns every seat, atomically create the booking, mark seats booked, mark H7 confirmed, record the payment result and emit E72. Return or notify confirmed only after commit.
  6. If the provider times out, retain P3 as unknown and reconcile the same attempt. If H7 has already released its seats, record successful payment plus refund-required work. Never try a new charge merely because a response was lost.

The provider's key retention and retry semantics are provider-specific. Our database retains P3 and its outcome beyond the interactive request so an expired provider retry window cannot silently become permission to charge again.

12Read and delivery path

Browsing may use a labeled stale seat map, while hold and booking-status reads must reflect authoritative ownership. This path separates catalog/map delivery, recovery after a pending payment, and waiting-list status.

  1. Customer A requests S99's map. The read service fetches immutable hall geometry and a versioned availability snapshot. It returns asOf=12:00:01 and quote version Q4; the browser can show a countdown only after receiving an authoritative hold deadline.
  2. On cache miss, the read service queries the show's authority or an explicitly permitted replica and constructs a new advisory map. A replica can lag; that is acceptable for browsing within the stated freshness budget, not for booking decisions. Suppress a cache stampede with one bounded refresh per show.
  3. After checkout returns pending, customer A requests H7's status or keeps a long poll open. Authenticate the customer’s ownership before returning seat or payment data. Route this status read to the authority when it must reflect the customer’s recent write; do not bounce between arbitrary lagging replicas.
  4. A committed E72 wakes a notification worker. It may deliver twice or late. The browser compares booking versions and fetches authoritative status if it observes a gap, rather than treating event arrival order as lifecycle order.
  5. Customer B asks about W8. The wait service returns waiting, admitted, canceled, expired or impossible with a server deadline; it need not promise an exact queue time because seat requirements differ. On admission, the admitted customer’s new hold attempt consumes the grant atomically.

Search and catalog results may use a separate index for text, location and date filtering. Removing a canceled show must also disable new holds at the authoritative write path immediately; waiting for the search index to refresh is not a correctness strategy.

13Correctness deep dive

Lock order and legal state transitions

Use fresh authority time after lock waits and one lock order: hold row, then its immutable seat IDs in ascending order, then the payment-attempt row when needed. Every lifecycle transaction re-reads the current row after acquiring the lock. The table describes durable effects within that transaction; the provider call happens outside it.

Event Preconditions checked under locks Atomic durable effect
Start checkout held, now < hold deadline, caller owns H7, all seats point to H7 H7 → processing, version + 1; store bounded processing deadline and unique P3
Hold expiry held, now ≥ hold deadline H7 → expired; release exactly seats still owned by H7; outbox event
Processing expiry processing, now ≥ processing deadline H7 → released; free its allocations; retain unresolved P3; schedule reconciliation
Verified success P3 belongs to H7; H7 processing and unexpired; every seat still owned by H7 H7 → confirmed; seats → booked; payment → succeeded; booking/outbox inserted
Late verified success H7 expired/released or no longer eligible Payment → succeeded; insert unique refund work; do not change seat owners
Verified decline/cancel Current eligible hold; external outcome known or cancellation policy applies Release allocations and record terminal state
transaction confirmSuccess(P3, verifiedOutcome):
  require verified SUCCEEDED outcome for stored provider/account/P3
  require amount and currency match the stored charge
  lock hold(P3.holdId); lock its seats in sorted order; lock P3
  if this success was already processed: return stored result
  record verified success
  if hold.state == PROCESSING and freshAuthorityTime() < hold.deadline
     and every seat.holdId == hold.id:
       create booking; mark all seats BOOKED; mark hold CONFIRMED
  else:
       insert refund_work(P3) ON CONFLICT DO NOTHING
  insert outbox event; commit

Confirmation wins the race

Race A: confirmation wins. At 12:05:59.900 success locks H7 first, checks its 12:06 deadline, and commits confirmed. The expiry worker later obtains the lock, sees confirmed, and does nothing. Customer B cannot claim seat 56 because it remains booked.

Expiry wins the race

Delayed cleanup does not extend a hold

An expiry worker crash can delay freeing seats but cannot validate an expired hold. A confirmation path itself rejects an overdue state; cleanup can then perform the guarded release. Refund work is also retried with the same refund identity, so recovery does not issue a new refund each time. This is not an atomic transaction across payment and inventory: it is an explicit policy for an unavoidable uncertain outcome.

Deadline time after acquiring locks

Read the deadline clock after acquiring the locks. PostgreSQL now()/CURRENT_TIMESTAMP describes transaction start, so a transaction begun at 12:05:59 can wait until 12:06:02 yet still report the earlier time. Use an actual current-time check such as clock_timestamp(), with a conservative clock-error policy, for new checkout/confirmation admission. Retrying an already-committed operation returns its saved result; it does not obtain another grace extension.

Verify successful payment meaning

state · hold-statesHold lifecycle and separate payment compensation

Expiry/release is terminal for the allocation. REFUND REQUIRED and REFUND RECORDED describe the associated payment recovery, not revived hold ownership.

Hold lifecycle and separate payment compensationExpiry/release is terminal for the allocation. REFUND REQUIRED and REFUND RECORDED describe the associated payment recovery, not revived hold ownership. held to processing: Checkout before deadline; persist P3; held to expired: Hold deadline / cancel; release seats; processing to confirmed: Success + eligible + owns every seat; processing to expired: Deadline / decline; release seats; expired to refund: Late success; keep new seat owners; refund to refunded: Verified idempotent refund outcomeCheckout before deadline;persist P3Hold deadline / cancel; releaseseatsSuccess + eligible + ownsevery seatDeadline / decline; releaseseatsLate success; keep new seatownersVerified idempotent refundoutcomeSTATEHELDSTATEPROCESSINGSTATECONFIRMEDSTATEEXPIRED / RELEASEDSTATEREFUND REQUIREDSTATEREFUND RECORDEDsync
Read each connection in order
  1. syncCheckout before deadline; persist P3HELD → PROCESSING
  2. syncHold deadline / cancel; release seatsHELD → EXPIRED / RELEASED
  3. syncSuccess + eligible + owns every seatPROCESSING → CONFIRMED
  4. syncDeadline / decline; release seatsPROCESSING → EXPIRED / RELEASED
  5. syncLate success; keep new seat ownersEXPIRED / RELEASED → REFUND REQUIRED
  6. syncVerified idempotent refund outcomeREFUND REQUIRED → REFUND RECORDED
sequence · expiry-winsRace: expiry wins before payment success

H7 remains released and H8 keeps its seats after the late callback. Under this design’s policy, the successful charge for H7 requires a refund.

Race: expiry wins before payment successH7 remains released and H8 keeps its seats after the late callback. Under this design’s policy, the successful charge for H7 requires a refund. expiry to db: Lock H7 at deadline; db to expiry: H7 processing / seats owned; expiry to db: Commit released + free seats; ben to db: Commit H8 for seats 56–57; callback to db: Lock H7; record P3 success; db to callback: H7 released; seat 56 now H8; callback to db: Insert unique refund work; no booking; db to callback: Commit; H8 remains ownerPARTICIPANTExpiry workerPARTICIPANTShow authority DBPARTICIPANTCompetingcustomerPARTICIPANTPayment handler1. Lock H7 at deadline2. H7 processing / seatsowned3. Commit released + freeseats4. Commit H8 for seats 56–575. Lock H7; record P3 success6. H7 released; seat 56 now H87. Insert unique refund work; no booking8. Commit; H8 remains ownersyncreturn
Read each connection in order
  1. syncLock H7 at deadlineExpiry worker → Show authority DB
  2. returnH7 processing / seats ownedShow authority DB → Expiry worker
  3. syncCommit released + free seatsExpiry worker → Show authority DB
  4. syncCommit H8 for seats 56–57Competing customer → Show authority DB
  5. syncLock H7; record P3 successPayment handler → Show authority DB
  6. returnH7 released; seat 56 now H8Show authority DB → Payment handler
  7. syncInsert unique refund work; no bookingPayment handler → Show authority DB
  8. returnCommit; H8 remains ownerShow authority DB → Payment handler

14Failure and recovery

Failure timeline User-visible result Surviving state and recovery
API dies after H7 commits, before reply customer A initially sees a timeout Retry k7 reads the committed result; no second hold
Provider accepts P3, network reply disappears Checkout stays pending Reconcile P3 by provider status/verified callback; never infer failure from timeout
Primary becomes isolated New writes pause or fail within deadline Only a safely promoted authority may accept holds; old owner is fenced before reopening
Dispatcher dies after sending E72 Duplicate notification is possible Browser reads booking version; outbox remains replayable
Expiry/coordinator restart Capacity and admission may be delayed Rebuild active deadlines and wait sequence from durable indexes, not volatile lists

During a 100× on-sale burst, bound active holds and database connections. The edge returns a waiting/admission response with backoff instead of allowing 50,000 transactions to wait on the same rows. Serve stale-but-labeled maps if necessary; never serve a synthetic successful hold. Separate provider timeout workers from interactive threads so a slow provider cannot occupy every booking slot.

The coordinator's epoch is stored in the show authority. A restarted coordinator advances it with an atomic compare-and-update, and grant creation requires that epoch to match. If the entire authority fails over, its database fencing policy still has to prevent two writable primaries; application epochs alone cannot repair split-brain storage. Restore tests must check both seats and waiting order. A restored seat table does not contain who arrived first, who canceled, or which payment remains unknown.

15Operations, security, and cost

Session ownership and payment verification

Authenticate guest/session ownership on every hold, status and cancellation request. Verify payment callbacks against their raw signed payload and deduplicate provider event IDs. Store provider tokens, not card details. Rate-limit seat hoarding by authenticated account/session and risk signals; an IP-only rule can punish a shared network without stopping a distributed bot. A signed admission grant is still checked for server-side consumption and expiry.

Hold, payment and waiting-list signals

Monitor hold p95, time waiting for locks, transaction aborts, active holds by age, provider-unknown backlog, refund age and wait-to-admission latency. Periodically query for impossible states, such as one seat allocated to two active booking records or a confirmed hold missing one requested seat. Alert the on-call operator when that seat-allocation rule is violated; a healthy HTTP success rate cannot prove inventory correctness.

Map egress and admission cost

The largest burst cost is often read egress and repeated map computation. A 99% cache hit reduces a 500 MB/s illustrative origin stream to 5 MB/s, but edge egress still exists. Three replicas turn 3.65 TB raw five-year inventory into at least 10.95 TB before indexes, logs and backups. Retaining every generated show-seat row forever may be unnecessary; archive completed shows under an explicit retrieval policy.

State rollout and timeout/interleaving tests

Roll out a new hold state by first deploying readers that understand it, then enabling writers for a small show cohort. Run crash tests at every commit/response boundary, replay duplicate callbacks and stop expiry for ten minutes. Recovery must preserve no-oversell even while availability temporarily degrades. Rehearse moving a shard: stop old-owner writes, copy the remaining changes, compare seat and hold counts, then route traffic to the new owner.

16Decision ledger and limitations

Decision Benefit Cost/consequence Revisit when
Relational, per-show transactions Direct all-seat atomicity One show's contention remains local Measured hot-show capacity needs a serialized command owner
Advisory cached maps Cheap burst browsing Users may lose a race after seeing green seats Product requires stronger reservation-like browsing semantics
Durable holds plus short locks Recoverable minutes-long checkout Expiry and payment transitions must be maintained Checkout policy changes to authorization-first or instant purchase
FIFO admission grants Fairness enforced at inventory Head-of-line blocking and idle grant windows Product explicitly permits bypass or lottery admission
One active write authority/show Clear allocation ownership Partition may pause booking A different globally coordinated transaction design is justified
Late-success refund policy Prevents overselling after release A customer may temporarily be charged without a booking Provider supports a suitable authorization/capture workflow

Do not promise that changing SQL to a key-value store eliminates the inventory conflict. It merely changes how the same atomic seat-set rule must be implemented. Optimistic versions can reduce lock overhead under low conflict but create retry storms for a popular seat. A per-show command queue can make ordering explicit, at the price of a new owner, queue latency and fenced failover. Choose it after measuring the simpler transaction path.

The residual limit is intentional: an exact scarce seat cannot be sold to unlimited simultaneous callers. We optimize useful work and clear outcomes, not the number of requests allowed to collide. The system also cannot make an external payment and local inventory commit instantaneously atomic; its documented compensation policy is part of the product.

17Interview closing

“I have separated browsing from reservation. A map is cheap and slightly stale; a hold is an authoritative all-or-nothing transaction over one show's seats. I start with one relational service, then cache map reads, add durable admission when contention grows, shard different shows and isolate recoverable background work. A stable request identity makes hold creation retry-safe, and checkout durably records its processing deadline and payment-attempt identity before calling the provider. Payment confirmation and expiry lock the same hold and seats, so only one eligible transition wins. A late successful charge becomes refund work without reclaiming seats released to another allocation.

“The main tradeoffs are stale maps, queue wait and temporarily pending payment outcomes. I accept write unavailability during an unsafe authority partition. One hot show is still limited by its inventory owner, so my next measurement is lock wait and completed hold transactions per second under overlapping seat sets, together with recovery behavior at the payment deadline.”

If the interviewer adds multi-show carts, the local atomic transaction no longer covers every seat set. Explain the new choice: reserve each show independently with a deadline and compensate partial holds, or use a distributed transaction system with the corresponding availability and operational cost. Do not simply add a cart service and retain the original atomicity claim. If they replace exact seats with interchangeable capacity, a guarded capacity counter can simplify the inventory model, but idempotency, holds and external payment uncertainty remain.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why not keep database row locks on selected seats for the five minutes a customer is allowed to pay?

Reveal a model answer

A row lock protects a short concurrent state transition. Holding it through human interaction consumes connections, prolongs contention and does not provide a recoverable checkout record. I commit a durable hold with an immutable seat set and server deadline, release the row locks, then use another short guarded transaction to confirm, cancel or expire it.

What the answer must demonstrate: Separate business lifetime from transaction lifetime.

Applied · Question 2

For one show, request A wants seats 54–56 and request B wants 56–57. How do you prevent both overselling seat 56 and partially reserving either request?

Reveal a model answer

Each transaction locks its requested seats in ascending order and validates the entire set before committing. If A commits first, seat 56 belongs to A’s hold. B then sees that conflict and rolls back its whole transaction, including any tentative change to seat 57. It returns conflict or a waiting option. If B wins first, the symmetric result applies; one request cannot retain only its uncontested seats.

What the answer must demonstrate: Trace actual row states and rollback scope.

Applied · Question 3

A payment-provider call times out while a seat hold approaches expiry. What state must exist before the call, and when may seats be released?

Reveal a model answer

Before calling the provider, I commit a PROCESSING hold state, a bounded processing deadline and one durable payment-attempt identity. A timeout leaves that attempt UNKNOWN; retries and reconciliation reuse it. Confirmation and expiry serialize on the hold and seats. The processing deadline may release inventory under the agreed policy even while the provider result is unknown, but a later successful charge must be compensated rather than taking those seats back.

What the answer must demonstrate: Timeout is an uncertain external outcome.

Follow-up · Question 4

Why does a FIFO waiting list not by itself guarantee fair booking?

Reveal a model answer

A newcomer could bypass the queue unless the hold endpoint consumes a current persisted admission grant. For the strict policy here, one unconsumed grant is active per contending show lane, issued to the oldest eligible ticket. The next ticket advances only after consumption, cancellation or expiry. Issuing several grants in order would not ensure redemption order, because a later customer’s network request might reach inventory first.

What the answer must demonstrate: Enforce queue policy where inventory is claimed.

Follow-up · Question 5

The reservation service and waiting service restart together. What is recovered?

Reveal a model answer

I reconstruct active holds from their durable state and deadlines, and waiting order from persisted ticket sequence and session expiry. The current coordinator resumes grants for each show; the database rejects grants from a replaced coordinator. Browser notifications may replay, so clients read authoritative status and tolerate duplicate messages.

What the answer must demonstrate: Name the missing state, not just replicas.

Foundation · Question 6

Why shard by show instead of movie?

Reveal a model answer

Seat contention is scoped to one show, and its seats should share a transaction domain. Movie partitioning places every showing of a blockbuster on one owner unnecessarily. Show partitioning distributes different performances while preserving local seat-set transactions.

What the answer must demonstrate: Distinguish distributed throughput from one-resource contention.

Applied · Question 7

Hold H7 expires and releases seats 54–56. Hold H8 then acquires 56–57 before H7’s payment reports success. What exact condition prevents the late callback from stealing seat 56?

Reveal a model answer

The success transaction locks H7 and its immutable seat set, then requires H7 to remain PROCESSING before its processing deadline and every seat to remain allocated to H7. Because H7 is released and seat 56 belongs to H8, that predicate is false. The handler records the verified payment result and inserts unique compensation work without restoring H7 or changing H8’s seats.

What the answer must demonstrate: State the atomic predicate and the losing outcome.

Follow-up · Question 8

Two waiting-list coordinators believe they own the same show. Where must the system reject a stale coordinator’s admission grants?

Reveal a model answer

The show authority stores one active coordinator epoch. Every grant-creation transaction checks that epoch while advancing the durable waiting sequence. A replacement obtains a newer epoch through an atomic update; an old coordinator cannot create a valid grant with the previous epoch. Hold creation then validates and consumes the persisted grant in the same inventory transaction. A signed token alone does not establish that it is current or unused.

What the answer must demonstrate: Identify the enforcing store and the remaining failure boundary.

Blank-page exercise · 45 minutes

Build the answer yourself

Design exact-seat booking with five-minute holds and a bounded payment-processing grace period. For one show, interleave requests for seats 54–56 and 56–57, then handle an unknown payment outcome at expiry while another customer waits. Prove no oversell and all-or-nothing allocation.

  • Model show-seat uniqueness and all-or-nothing orders.
  • Distinguish advisory browsing, durable holds, and row locks.
  • Trace the two overlapping transactions.
  • Resolve payment/expiry using explicit versions and deadlines.
  • Recover fair waiting and unknown payment state.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a ticket-booking serviceWhat persists during payment?Recall first, then reveal

A durable expiring hold; database row locks have already been released.

Lock briefly, hold durably.

Return to lesson
Design a ticket-booking serviceWhat does all-or-nothing mean?Recall first, then reveal

The order receives every requested seat or none. A conflict rolls back the complete requested set rather than leaving an unnoticed partial reservation.

One order, one atomic seat set.

Return to lesson
Design a ticket-booking serviceWhat does a payment timeout mean?Recall first, then reveal

The result is unknown; reconcile the same attempt before taking a new financial action.

Unknown is not declined.

Return to lesson

Final revision

Summary and interview notes

A seat map suggests availability; only a committed hold reserves the whole requested seat set. Short transactions control confirmation, cancellation and expiry. If payment succeeds after the seats were released, record refund work without taking seats from another customer.

Remember these points

  • A durable hold lasts minutes, while its database locks exist only for short state transitions.
  • Claim the request identity before allocating seats so concurrent identical requests recover one result.
  • Checkout records one payment attempt and bounded processing deadline before contacting the provider.
  • Confirmation and expiry use the same lock order, fresh deadline time and current ownership of every requested seat.
  • Strict waiting priority requires enforcement at grant redemption; FIFO grant issuance alone does not preserve allocation order.

Interview tips

  • Interleave overlapping requests for seats 54–56 and 56–57 and show the losing transaction leaves no partial hold.
  • Pause the provider response past expiry, allocate a released seat elsewhere, then show why late success creates only refund work.
  • Separate map QPS, admitted hold capacity and strict waiting-lane throughput.

Important qualifications

  • Webhook authenticity is separate from matching the stored provider attempt, amount, currency and successful outcome.
  • Provider idempotency retention is finite and does not replace durable business attempt history.
  • A single show remains a contention domain; sharding different shows or adding replicas does not remove that limit.

Technical references

Practice marks stay in this browser.