System-design interview · Core interviews
Design a ticket-booking service
Design all-or-nothing seat allocation, durable holds, fair admission and payment recovery; preserve inventory ownership when expiry, retries and late provider outcomes race.
You will learn to
- Demonstrate all-or-nothing reservation of overlapping seat sets.
- Separate a five-minute business hold from a short database lock.
- Recover hold expiry, waiting order, and ambiguous payment outcomes.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Databases, data models, and ACID transactions · Quorums, consensus, leases, and fencing · Message queues, event logs, delivery guarantees, and backpressure · Caching: cache hits, misses, write policies and invalidation
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Problem and scope
A ticket-booking service sells scarce seat inventory while allowing browsing, temporary holds and payment. Its core invariant is that one show-seat has at most one current allocation and a requested seat set commits entirely or not at all. For example, customer A requests seats 54–56 while customer B requests 56–57 for show S99. Both maps may show seat 56 as available, but the authoritative booking transaction must allow only one competing seat set to succeed.
Scope the product to exact seat selection with all-or-nothing requests of up to ten seats, a five-minute hold, and an explicit bounded payment-processing grace period. Include city/movie/cinema/show discovery, seat maps, payment confirmation, cancellation and durable waiting. Resale, cross-show carts and dynamic auction pricing are excluded. This is an interview design, not a claim about Ticketmaster’s internal implementation.
The payment-processing grace period is another server deadline: it bounds how long an eligible hold remains allocated while an initiated payment is being resolved. It does not make a payment timeout a decline or let repeated checkout requests extend the hold indefinitely.
02Functional requirements
- Search and browse: Search by city/postal code or coordinates/radius, keyword, date and showtime; sort/paginate and offer spelling help where useful. Browse movie → cinema → hall → show → seat map, including seat class and price.
- Hold an exact seat set: Success reserves every requested seat under one hold; conflict reserves none. A free-looking map seat remains advisory until allocation commits.
- Pay and recover status: Pay for a hold, observe pending/confirmed/expired, and retrieve the booking after a lost response. A payment timeout is not a definite decline.
- Cancel a hold: Cancel when product rules permit.
- Wait for admission: Join, inspect or leave a durable per-show waiting list while capacity is temporarily held; notify the browser when admission or booking status changes.
Guest ownership and waiting limits
Guest checkout uses a server-issued unguessable session credential, not a caller-supplied name. A returning guest recovers the same hold through that credential. Waiting ends on cancellation, a one-hour session limit or a provably impossible request. Zero free seats alone does not prove impossibility because existing holds can expire.
Strict FIFO admission protocol
Admission controls who may attempt to create a hold when too many customers are competing. An admission grant is temporary permission to make that attempt; it does not itself reserve seats. The hold transaction must validate and consume the grant while allocating the requested seat set.
First-in-first-out (FIFO) admission serves waiting customers in their persisted order. A show lane is the queue of requests competing for the same show inventory; this strict version allows only one customer in that lane to hold an unused admission grant at a time.
- Select the oldest eligible ticket: Persist waiting order and allow at most one unconsumed admission grant per contending show lane.
- Issue a short grant: The hold endpoint must validate that active grant.
- Consume and allocate atomically: The hold transaction consumes the grant with seat allocation. Later grant expiry cannot release a hold already created from it.
- Advance the lane: Admit the next ticket only after the prior grant is consumed, canceled or expires.
| Admission choice | Consequence |
|---|---|
| Strict FIFO | A six-adjacent-seat request can block smaller requests when only scattered singles remain. |
| Configured bypass | Improves utilization but changes the fairness promise. |
| Multiple simultaneous grants | FIFO issuance does not ensure FIFO redemption or strict seat-allocation priority. |
A wider admission window is an explicit higher-throughput fairness alternative. Explain this choice rather than promising both strict FIFO and maximum utilization.
03Non-functional requirements
- Read latency: Catalog/seat-map p95 below 200 ms; the advisory seat map may be stale for up to two seconds.
- Booking latency: Admitted hold creation p95 below 500 ms; status updates within two seconds of a committed transition. Requests past their deadline fail clearly rather than queue indefinitely.
- Availability: 99.95% monthly availability for admitted booking requests. Inventory safety takes priority over accepting writes during an authority partition.
- Durability: A successful hold survives one database-node failure. Use three replicas across failure domains and an explicitly configured durable commit policy; replica count alone does not establish a recovery point objective (RPO), the maximum acknowledged-data loss allowed after a failure.
- Regional recovery: Assume a tested 15-minute restoration objective and disclose any asynchronous cross-region recovery-point gap.
- Retention and payment security: Size booking/payment audit records for an assumed five years, subject to business policy. Keep payment credentials at the provider. Expired sessions and wait tickets need much shorter retention.
Inventory and payment invariants
| Invariant | Required behavior |
|---|---|
| Exclusive seat allocation | (showId,seatId) has at most one current allocation. |
| Atomic seat set | An order gets every requested seat or none. |
| Terminal hold state | An expired/released hold cannot steal seats back. |
| Stable provider identity | One logical payment attempt uses one provider identity. |
| Authoritative confirmation | Neither a browser redirect nor a stale map approves a reservation. |
These are interview assumptions to negotiate, not measured product facts. A latency/availability target must not weaken the seat-allocation contract.
04Capacity estimates
Workload assumptions and arithmetic
Assume 3 billion page views and 10 million ticket sales per month. At 30 days/month, these are about 1,157 views/s and 3.86 tickets/s, not orders/s. At two tickets/order, the average is only 1.93 orders/s. That modest average must not justify an unbounded checkout endpoint for a popular on-sale event.
Worked estimates
| Quantity | Calculation | Consequence |
|---|---|---|
| Daily show-seat inventory | 500 cities × 10 cinemas × 2,000 seats × 2 shows | 20 million rows/day |
| Raw inventory storage | 20M × 100 bytes × 1,825 days | 3.65 TB before indexes and replicas |
| Illustrative burst | 100,000 shoppers opening one show in 10 s | 10,000 seat-map reads/s |
| Map payload | 10,000/s × 50 KB | 500 MB/s if every response reaches origin |
| Hold admission budget | 500 successful/failed attempts/s × 20 ms mean service | About ten concurrent active transactions before lock waiting |
Capacity implications and limits
The last number is an initial benchmark target, not a claimed database capability. Contention on the same seat can dominate even when CPU is idle. Cache catalog and advisory maps; keep admission proportional to measured lock service capacity. At a five-minute hold lifetime, 500 admitted holds/s could create 150,000 active holds, although the available seats of one show impose a much smaller practical ceiling. Across many shows, index (state, deadline) for incremental expiry scans.
Store physical seat geometry and movie/hall metadata once. Repeating a hall layout inside every show-seat row inflates both cache and storage estimates. Price snapshots and the exact selected seats do belong with the order so later catalog edits cannot change an already-agreed charge.
The 500-attempt/s figure is an authority load-test budget, not the achieved rate of the strict FIFO waiting lane. One outstanding grant may be limited by client round-trip or its expiry interval, and a difficult head request can leave inventory idle. Measure that separately; obtaining higher throughput by issuing many concurrent grants changes the fairness contract.
05APIs and contracts
Interface contracts
| API | Successful result | Important failure |
|---|---|---|
GET /shows/S99/seats |
Layout, advisory states, asOf, map version |
A map never constitutes a hold |
POST /shows/S99/holds |
201 {holdId:H7, seats:[54,55,56], expiresAt, version:1} |
409 seat_conflict, 429 admission_required |
POST /holds/H7/checkout |
202 {paymentAttempt:P3,state:processing,deadline} |
409 expired or request-payload mismatch |
GET /bookings/by-hold/H7 |
Pending, confirmed, expired or refund-recovery status | Unauthorized ownership returns no private details |
POST /shows/S99/wait-tickets |
Ticket W8 and current state | Request exceeds show capacity or session limit |
Validation and response semantics
The hold request contains Idempotency-Key: k7, seat IDs, an expected quoted price version and, when required, an admission grant. The server obtains session identity from authentication. Persist (showId, sessionId, key) with a canonical payload fingerprint and result in the same transaction as H7. Reusing k7 with another seat set is a conflict; retrying the same request returns H7 even after a response is lost. Specify retention long enough to cover the user retry window, and require explicit status recovery after that window.
The booking service must recover payment outcomes independently of the browser. A webhook is a provider-to-service notification delivered to a registered endpoint; it may arrive after the checkout request has timed out. The stored payment attempt connects that notification to the original hold.
The client supplies a new checkout request key, but the service creates and durably retains provider payment identity P3. A retried checkout reuses P3. Webhooks carry provider event IDs and require signature verification; duplicated or out-of-order callbacks must be reconciled against the durable attempt state. Search pagination uses an opaque cursor carrying filters and sort position, not an ever-increasing offset through a changing catalog.
06Data model and access patterns
ShowSeat tracks the allocation of one physical seat for one show; Hold groups the requested seat set, while PaymentAttempt tracks the separate external charge. RequestResult saves the response needed for retries. The outbox stores notification work in the same transaction as a booking change, so later delivery can recover from a crash.
| Record | Example key and fields | Authority/query |
|---|---|---|
| ShowSeat | (S99,56): state=held,holdId=H7,version=1 |
Booking authority; exact seat-set lookup |
| Hold | (S99,H7): session=sessionA,state=held,deadline=12:05,version=1 |
Authority for lifecycle and ownership |
| HoldSeat | (S99,H7,56) |
Immutable requested set, also 54 and 55 |
| PaymentAttempt | (S99,P3): hold=H7,providerId, outcome=unknown |
External outcome and reconciliation history |
| WaitTicket | (S99,108): W8,requested=2,state=waiting,expiresAt=13:00 |
Durable FIFO order |
| AdmissionGrant | (S99,G9): ticket=W8,consumed=false,deadline |
Required booking-path permission |
| RequestResult / Outbox | (S99,session,key) / (S99,eventId) |
Replay response / recover notifications |
Shard by ShowID, so seats, hold, admission and outbox for S99 share one relational transaction domain. Sharding by MovieID would put every showing of a blockbuster together without improving the seat invariant. City, Cinema, Hall, Seat, Movie and Show form the catalog relations; their materialized search index and seat-map cache are derived stores.
Create indexes on active holds by deadline, bookings by owner/show, waiting tickets by show/sequence and outbox entries by dispatch state. To validate H7, lock its hold row, load its immutable seat set, then lock seats in ascending ID order. Use the same order for every hold lifecycle operation. New hold creation locks requested seats in ascending order. Database constraints prevent duplicate request IDs; transaction logic enforces that all seats move together. An index cannot by itself encode every multi-row lifecycle rule.
For expiry, use the durable ordered deadline index to find due holds and an exact hold-ID lookup for cancellation/removal. A linked list ordered only by creation time is valid for expiry scanning only when creation order also preserves deadline order. Payment grace periods can change that order, so deadline indexing is the safer baseline. After worker restart, rebuild ready batches from stored active holds and waiting tickets rather than trusting a replicated in-memory list as the only recovery record.
For creation, a database unique constraint claims the request identity inside the same transaction as all seat effects. A conflicting duplicate rolls back any tentative work and reads the saved payload/result in a fresh transaction if required by the isolation mode. Checking for a missing request row before locking seats is not enough: both concurrent attempts can see absence, and the loser could otherwise report conflict after the winner already created its hold.
07Basic working design
One booking process and SQL authority
The first version has a browser, one booking process, a relational database and an external payment provider. Catalog reads, hold creation, expiry scans and payment reconciliation run in this process. This is enough to prove business behavior before introducing independent queues or caches.
Atomic all-seat hold creation
Customer A submits k7 for seats 54–56. The service authenticates the customer session, begins a transaction, claims the unique request-result key, and locks all three ShowSeat rows. A concurrent duplicate waits on that key and returns the committed matching result before attempting seat allocation; it does not report a seat conflict against its own successful first attempt. It verifies every seat is free and the quote is still valid. It inserts H7 and its HoldSeat rows, changes all three allocations, inserts the replay result, and commits. Only then does it return 201 H7. A crash before commit leaves no hold; a crash after commit but before response leaves H7 recoverable by k7.
Conflicts and payment intent
Customer B's request for 56–57 cannot partially succeed. If customer A owns the conflicting seat when customer B obtains the locks, the competing transaction rolls back. The customer can request alternatives or wait. The service does not lock seats while calling the provider: it first commits a bounded processing state and P3, then makes the network call.
Deployable baseline and limits
This baseline can run as a complete service: a periodic loop expires holds and reconciles payments from saved pending work, including after a restart. Its limitations are throughput, process availability and operational isolation; its basic inventory invariant is already valid.
One transaction changes the entire seat set. Payment remains an external boundary.
Read each connection in order
- sync1. Hold k7 / seats 54–56Customer browser → Booking application
- sync2. Commit all-seat H7 transactionBooking application → Relational booking database
- sync3. Durable H7 resultRelational booking database → Booking application
- sync4. Return hold and deadlineBooking application → Customer browser
- sync5. Pay only after P3 is durableBooking application → Payment provider
08Find the baseline flaws
| Bottleneck / counterexample | Evidence and design consequence |
|---|---|
| On-sale read and hold stampede | First replay the on-sale workload: 10,000 map reads/s and a stampede of hold attempts all reach one process and one database. At 50 KB/map the process may try to emit 500 MB/s before TLS and query overhead. A popular seat causes waiting transactions that occupy connections; adding more web threads merely grows the waiting room inside the database. The performance bottleneck is partly repeated browsing work and partly serialized inventory contention. They need different changes. |
| Late payment after hold expiry | Then inject a correctness failure into a tempting implementation. At 12:04:58 checkout calls the provider without first changing H7's durable state. At 12:05:00 expiry frees seats 54–56. Customer B then buys 56. At 12:05:01 the provider reports success for customer A. Code that blindly marks H7 confirmed oversells. Adding a cache or more replicas does not repair that interleaving. |
| Required durable payment state | The correct baseline already needs a transaction that records processing, its deadline and P3 before the provider call; final confirmation must re-check both hold state and allocation ownership under locks. If expiry wins a later race, a successful external charge must create refund work, the compensating action for a charge without a booking. Test this by pausing the provider response, running expiry and then releasing the response. Record actual row values after each commit. This test defines the state machine that the larger design must preserve. |
09Improve the design, step by step
Change 1 — move advisory reads off the booking database. The trigger is the 500 MB/s map burst. Cache catalog and short-lived map snapshots behind the read service, while serving static geometry through an edge cache. At a 99% map hit rate, 10,000 requests/s become roughly 100 origin reads/s. The costs are cache memory, invalidation traffic and stale displays. The new failure is a convincing-looking stale map, so every hold still validates authority. Direct database reads remain preferable for a small deployment where freshness and simplicity outweigh cache savings.
Change 2 — enforce durable admission. The trigger is lock queues exceeding the 500 ms hold objective. The wait coordinator grants a limited number of short-lived tokens in persisted sequence order; the hold transaction consumes the token. This bounds work reaching scarce seats and prevents newcomers bypassing waiting users. The cost is extra waiting latency, durable coordinator state and idle capacity when grants expire unused. The new risk is a restarted coordinator issuing stale grants: the database validates the active coordinator’s ownership version, called its epoch, when creating them. A simple reject-and-retry limit is cheaper when FIFO fairness is not a product promise.
Change 3 — partition by show and replicate each authority. The trigger is aggregate inventory exceeding one database's measured write or storage capacity. A routing directory maps S99 to shard 12; all of its seat sets remain local transactions. Different shows gain parallel capacity; one hot show does not. Costs include rebalancing, replica lag and shard-aware operations. Stale routing can send requests to the old owner, so the handoff must prevent that shard from writing before the new owner accepts requests. A larger single database is preferable until this complexity buys measured headroom.
Change 4 — extract durable background work. The trigger is payment timeouts and notification retries competing with interactive holds. Commit outbox records beside booking changes; dispatchers feed wait admission, map invalidation and browser notifications. Expiry and payment reconcilers use indexed durable work. This improves isolation and recovery, at the cost of duplicate events, lag and extra workers. Consumers deduplicate event IDs and fetch authoritative status. An in-process worker is still suitable while it has the same persistent work contract and adequate isolation.
10Detailed architecture
Read and booking request paths
The browser enters through an authentication/admission edge. Browsing goes to the catalog/map service and derived cache; booking writes go through the show router to the booking authority. The router consults a show-to-shard directory, which is configuration, not inventory. The final diagram separates these paths because a cached “available” response and a committed hold have different guarantees.
Show-local source of truth
Inside a show shard, the booking service and relational primary own ShowSeat, Hold, PaymentAttempt, waiting/admission state and outbox records. Replicas implement the promised durable commit and recovery policy. The application acknowledges success only after that policy completes. Other regions may serve catalog reads, but a single active authority accepts seat writes for S99. Promoting another owner requires fencing the previous writer through the storage failover mechanism; a DNS change alone is insufficient.
Expiry, reconciliation and notifications
Expiry/reconciliation workers and outbox dispatchers sit outside the synchronous response path. Each worker checks the saved hold or payment state in a transaction before changing it. The payment provider is an external transaction boundary. Its request and our local booking update cannot be wrapped in one database transaction. Browser notifications report changes; they never create allocations. The wait coordinator creates grants in the same authority that consumes them. If it cannot safely issue a grant, customers stay waiting while existing valid holds remain readable.
Hot-show transaction bottleneck
This architecture has an intentionally visible bottleneck: the transactional owner of one very popular show. Admission makes that limit survivable; replication and sharding do not abolish it.
Seat-map caches cannot grant inventory. Every hold, grant and lifecycle transition reaches the same show authority.
Read each connection in order
- sync1. Authenticated browse / holdCustomer browsers → Identity / admission edge
- sync2a. Browse S99Identity / admission edge → Catalog and map service
- syncRead advisory snapshotCatalog and map service → Advisory map / catalog cache
- syncRefresh map on cache missCatalog and map service → Show router
- sync2b. Hold k7 with admission grantIdentity / admission edge → Show router
- sync3. Resolve S99 ownerShow router → Show → shard directory
- sync4. Route to authoritative shardShow router → Booking authority
- sync5. Atomic seats + H7 + replayBooking authority → Show-shard primary Holds / seats / outbox
- replicationReplicate durable commitShow-shard primary Holds / seats / outbox → Durable shard replicas
- sync6. Submit stable payment P3Booking authority → Payment provider
- sync7. Verified payment outcomePayment provider → Booking authority
- syncGuard epoch; create active FIFO grantWait / grant coordinator → Show-shard primary Holds / seats / outbox
- syncGuarded expiry / unknown outcomesExpiry / payment reconciler → Show-shard primary Holds / seats / outbox
- syncReconcile P3 / refund identityExpiry / payment reconciler → Payment provider
- async8. Read committed outboxShow-shard primary Holds / seats / outbox → Outbox / status dispatcher
- asyncInvalidate snapshot versionOutbox / status dispatcher → Advisory map / catalog cache
- asyncCapacity / queue eventOutbox / status dispatcher → Wait / grant coordinator
- asyncNotify; client fetches statusOutbox / status dispatcher → Customer browsers
11Write path and acknowledgement
A hold and a payment attempt are separate durable records. The example defines show S99, customer session sessionA, request k7, seats 54–56, hold H7, admission grant G9 and provider attempt P3 before tracing their commits.
- Customer A sends k7, quote version Q4, G9 and seats 54–56. The edge authenticates sessionA and routes S99 to shard 12; request limits stop repeated speculative holds.
- The transaction first claims the unique
(S99,sessionA,k7)request row, serializing concurrent duplicates before grant consumption or seat locks. An existing identical request returns H7; a changed payload conflicts. It verifies and consumes G9 when waiting is active, locks seats in order and validates every allocation and quoted price. - It writes H7 version 1, all seat allocations, the replay result and outbox event E71. After the durability policy succeeds, return the server deadline. A lost response is recovered with the same k7.
- Checkout locks H7 and its seats and first returns any matching saved checkout attempt. For a new checkout, if H7 is held and unexpired according to fresh authority time after lock waits, change it to processing version 2 with deadline 12:06, create P3 and commit. The provider call uses P3's stable idempotency identity and the stored amount/currency.
- A verified success follows the common order: H7, its seats in ascending order, then P3. If H7 is still valid processing and owns every seat, atomically create the booking, mark seats booked, mark H7 confirmed, record the payment result and emit E72. Return or notify confirmed only after commit.
- If the provider times out, retain P3 as unknown and reconcile the same attempt. If H7 has already released its seats, record successful payment plus refund-required work. Never try a new charge merely because a response was lost.
The provider's key retention and retry semantics are provider-specific. Our database retains P3 and its outcome beyond the interactive request so an expired provider retry window cannot silently become permission to charge again.
12Read and delivery path
Browsing may use a labeled stale seat map, while hold and booking-status reads must reflect authoritative ownership. This path separates catalog/map delivery, recovery after a pending payment, and waiting-list status.
- Customer A requests S99's map. The read service fetches immutable hall geometry and a versioned availability snapshot. It returns
asOf=12:00:01and quote version Q4; the browser can show a countdown only after receiving an authoritative hold deadline. - On cache miss, the read service queries the show's authority or an explicitly permitted replica and constructs a new advisory map. A replica can lag; that is acceptable for browsing within the stated freshness budget, not for booking decisions. Suppress a cache stampede with one bounded refresh per show.
- After checkout returns pending, customer A requests H7's status or keeps a long poll open. Authenticate the customer’s ownership before returning seat or payment data. Route this status read to the authority when it must reflect the customer’s recent write; do not bounce between arbitrary lagging replicas.
- A committed E72 wakes a notification worker. It may deliver twice or late. The browser compares booking versions and fetches authoritative status if it observes a gap, rather than treating event arrival order as lifecycle order.
- Customer B asks about W8. The wait service returns waiting, admitted, canceled, expired or impossible with a server deadline; it need not promise an exact queue time because seat requirements differ. On admission, the admitted customer’s new hold attempt consumes the grant atomically.
Search and catalog results may use a separate index for text, location and date filtering. Removing a canceled show must also disable new holds at the authoritative write path immediately; waiting for the search index to refresh is not a correctness strategy.
13Correctness deep dive
Lock order and legal state transitions
Use fresh authority time after lock waits and one lock order: hold row, then its immutable seat IDs in ascending order, then the payment-attempt row when needed. Every lifecycle transaction re-reads the current row after acquiring the lock. The table describes durable effects within that transaction; the provider call happens outside it.
| Event | Preconditions checked under locks | Atomic durable effect |
|---|---|---|
| Start checkout | held, now < hold deadline, caller owns H7, all seats point to H7 |
H7 → processing, version + 1; store bounded processing deadline and unique P3 |
| Hold expiry | held, now ≥ hold deadline |
H7 → expired; release exactly seats still owned by H7; outbox event |
| Processing expiry | processing, now ≥ processing deadline |
H7 → released; free its allocations; retain unresolved P3; schedule reconciliation |
| Verified success | P3 belongs to H7; H7 processing and unexpired; every seat still owned by H7 | H7 → confirmed; seats → booked; payment → succeeded; booking/outbox inserted |
| Late verified success | H7 expired/released or no longer eligible | Payment → succeeded; insert unique refund work; do not change seat owners |
| Verified decline/cancel | Current eligible hold; external outcome known or cancellation policy applies | Release allocations and record terminal state |
transaction confirmSuccess(P3, verifiedOutcome):
require verified SUCCEEDED outcome for stored provider/account/P3
require amount and currency match the stored charge
lock hold(P3.holdId); lock its seats in sorted order; lock P3
if this success was already processed: return stored result
record verified success
if hold.state == PROCESSING and freshAuthorityTime() < hold.deadline
and every seat.holdId == hold.id:
create booking; mark all seats BOOKED; mark hold CONFIRMED
else:
insert refund_work(P3) ON CONFLICT DO NOTHING
insert outbox event; commit
Confirmation wins the race
Race A: confirmation wins. At 12:05:59.900 success locks H7 first, checks its 12:06 deadline, and commits confirmed. The expiry worker later obtains the lock, sees confirmed, and does nothing. Customer B cannot claim seat 56 because it remains booked.
Expiry wins the race
Delayed cleanup does not extend a hold
An expiry worker crash can delay freeing seats but cannot validate an expired hold. A confirmation path itself rejects an overdue state; cleanup can then perform the guarded release. Refund work is also retried with the same refund identity, so recovery does not issue a new refund each time. This is not an atomic transaction across payment and inventory: it is an explicit policy for an unavoidable uncertain outcome.
Deadline time after acquiring locks
Read the deadline clock after acquiring the locks. PostgreSQL now()/CURRENT_TIMESTAMP describes transaction start, so a transaction begun at 12:05:59 can wait until 12:06:02 yet still report the earlier time. Use an actual current-time check such as clock_timestamp(), with a conservative clock-error policy, for new checkout/confirmation admission. Retrying an already-committed operation returns its saved result; it does not obtain another grace extension.
Verify successful payment meaning
Expiry/release is terminal for the allocation. REFUND REQUIRED and REFUND RECORDED describe the associated payment recovery, not revived hold ownership.
Read each connection in order
- syncCheckout before deadline; persist P3HELD → PROCESSING
- syncHold deadline / cancel; release seatsHELD → EXPIRED / RELEASED
- syncSuccess + eligible + owns every seatPROCESSING → CONFIRMED
- syncDeadline / decline; release seatsPROCESSING → EXPIRED / RELEASED
- syncLate success; keep new seat ownersEXPIRED / RELEASED → REFUND REQUIRED
- syncVerified idempotent refund outcomeREFUND REQUIRED → REFUND RECORDED
H7 remains released and H8 keeps its seats after the late callback. Under this design’s policy, the successful charge for H7 requires a refund.
Read each connection in order
- syncLock H7 at deadlineExpiry worker → Show authority DB
- returnH7 processing / seats ownedShow authority DB → Expiry worker
- syncCommit released + free seatsExpiry worker → Show authority DB
- syncCommit H8 for seats 56–57Competing customer → Show authority DB
- syncLock H7; record P3 successPayment handler → Show authority DB
- returnH7 released; seat 56 now H8Show authority DB → Payment handler
- syncInsert unique refund work; no bookingPayment handler → Show authority DB
- returnCommit; H8 remains ownerShow authority DB → Payment handler
14Failure and recovery
| Failure timeline | User-visible result | Surviving state and recovery |
|---|---|---|
| API dies after H7 commits, before reply | customer A initially sees a timeout | Retry k7 reads the committed result; no second hold |
| Provider accepts P3, network reply disappears | Checkout stays pending | Reconcile P3 by provider status/verified callback; never infer failure from timeout |
| Primary becomes isolated | New writes pause or fail within deadline | Only a safely promoted authority may accept holds; old owner is fenced before reopening |
| Dispatcher dies after sending E72 | Duplicate notification is possible | Browser reads booking version; outbox remains replayable |
| Expiry/coordinator restart | Capacity and admission may be delayed | Rebuild active deadlines and wait sequence from durable indexes, not volatile lists |
During a 100× on-sale burst, bound active holds and database connections. The edge returns a waiting/admission response with backoff instead of allowing 50,000 transactions to wait on the same rows. Serve stale-but-labeled maps if necessary; never serve a synthetic successful hold. Separate provider timeout workers from interactive threads so a slow provider cannot occupy every booking slot.
The coordinator's epoch is stored in the show authority. A restarted coordinator advances it with an atomic compare-and-update, and grant creation requires that epoch to match. If the entire authority fails over, its database fencing policy still has to prevent two writable primaries; application epochs alone cannot repair split-brain storage. Restore tests must check both seats and waiting order. A restored seat table does not contain who arrived first, who canceled, or which payment remains unknown.
15Operations, security, and cost
Session ownership and payment verification
Authenticate guest/session ownership on every hold, status and cancellation request. Verify payment callbacks against their raw signed payload and deduplicate provider event IDs. Store provider tokens, not card details. Rate-limit seat hoarding by authenticated account/session and risk signals; an IP-only rule can punish a shared network without stopping a distributed bot. A signed admission grant is still checked for server-side consumption and expiry.
Hold, payment and waiting-list signals
Monitor hold p95, time waiting for locks, transaction aborts, active holds by age, provider-unknown backlog, refund age and wait-to-admission latency. Periodically query for impossible states, such as one seat allocated to two active booking records or a confirmed hold missing one requested seat. Alert the on-call operator when that seat-allocation rule is violated; a healthy HTTP success rate cannot prove inventory correctness.
Map egress and admission cost
The largest burst cost is often read egress and repeated map computation. A 99% cache hit reduces a 500 MB/s illustrative origin stream to 5 MB/s, but edge egress still exists. Three replicas turn 3.65 TB raw five-year inventory into at least 10.95 TB before indexes, logs and backups. Retaining every generated show-seat row forever may be unnecessary; archive completed shows under an explicit retrieval policy.
State rollout and timeout/interleaving tests
Roll out a new hold state by first deploying readers that understand it, then enabling writers for a small show cohort. Run crash tests at every commit/response boundary, replay duplicate callbacks and stop expiry for ten minutes. Recovery must preserve no-oversell even while availability temporarily degrades. Rehearse moving a shard: stop old-owner writes, copy the remaining changes, compare seat and hold counts, then route traffic to the new owner.
16Decision ledger and limitations
| Decision | Benefit | Cost/consequence | Revisit when |
|---|---|---|---|
| Relational, per-show transactions | Direct all-seat atomicity | One show's contention remains local | Measured hot-show capacity needs a serialized command owner |
| Advisory cached maps | Cheap burst browsing | Users may lose a race after seeing green seats | Product requires stronger reservation-like browsing semantics |
| Durable holds plus short locks | Recoverable minutes-long checkout | Expiry and payment transitions must be maintained | Checkout policy changes to authorization-first or instant purchase |
| FIFO admission grants | Fairness enforced at inventory | Head-of-line blocking and idle grant windows | Product explicitly permits bypass or lottery admission |
| One active write authority/show | Clear allocation ownership | Partition may pause booking | A different globally coordinated transaction design is justified |
| Late-success refund policy | Prevents overselling after release | A customer may temporarily be charged without a booking | Provider supports a suitable authorization/capture workflow |
Do not promise that changing SQL to a key-value store eliminates the inventory conflict. It merely changes how the same atomic seat-set rule must be implemented. Optimistic versions can reduce lock overhead under low conflict but create retry storms for a popular seat. A per-show command queue can make ordering explicit, at the price of a new owner, queue latency and fenced failover. Choose it after measuring the simpler transaction path.
The residual limit is intentional: an exact scarce seat cannot be sold to unlimited simultaneous callers. We optimize useful work and clear outcomes, not the number of requests allowed to collide. The system also cannot make an external payment and local inventory commit instantaneously atomic; its documented compensation policy is part of the product.
17Interview closing
“I have separated browsing from reservation. A map is cheap and slightly stale; a hold is an authoritative all-or-nothing transaction over one show's seats. I start with one relational service, then cache map reads, add durable admission when contention grows, shard different shows and isolate recoverable background work. A stable request identity makes hold creation retry-safe, and checkout durably records its processing deadline and payment-attempt identity before calling the provider. Payment confirmation and expiry lock the same hold and seats, so only one eligible transition wins. A late successful charge becomes refund work without reclaiming seats released to another allocation.
“The main tradeoffs are stale maps, queue wait and temporarily pending payment outcomes. I accept write unavailability during an unsafe authority partition. One hot show is still limited by its inventory owner, so my next measurement is lock wait and completed hold transactions per second under overlapping seat sets, together with recovery behavior at the payment deadline.”
If the interviewer adds multi-show carts, the local atomic transaction no longer covers every seat set. Explain the new choice: reserve each show independently with a deadline and compensate partial holds, or use a distributed transaction system with the corresponding availability and operational cost. Do not simply add a cart service and retain the original atomicity claim. If they replace exact seats with interchangeable capacity, a guarded capacity counter can simplify the inventory model, but idempotency, holds and external payment uncertainty remain.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Why not keep database row locks on selected seats for the five minutes a customer is allowed to pay?
Reveal a model answer
A row lock protects a short concurrent state transition. Holding it through human interaction consumes connections, prolongs contention and does not provide a recoverable checkout record. I commit a durable hold with an immutable seat set and server deadline, release the row locks, then use another short guarded transaction to confirm, cancel or expire it.
Interviewer follow-up
What if the expiry daemon stops?
Reveal the follow-up answer
Use the stored server deadline when validating a hold, not only daemon cleanup. The worker restores prompt availability and notifications, but a late cleanup must not make an expired hold valid. Check fresh authority time after acquiring locks; a transaction-start timestamp can be stale after waiting.
What the answer must demonstrate: Separate business lifetime from transaction lifetime.
For one show, request A wants seats 54–56 and request B wants 56–57. How do you prevent both overselling seat 56 and partially reserving either request?
Reveal a model answer
Each transaction locks its requested seats in ascending order and validates the entire set before committing. If A commits first, seat 56 belongs to A’s hold. B then sees that conflict and rolls back its whole transaction, including any tentative change to seat 57. It returns conflict or a waiting option. If B wins first, the symmetric result applies; one request cannot retain only its uncontested seats.
Interviewer follow-up
What prevents deadlocks when sets overlap differently?
Reveal the follow-up answer
Acquire rows in a consistent order and keep transactions short. Still handle deadlock/serialization aborts with bounded retries; ordering reduces risk but does not justify ignoring database errors.
What the answer must demonstrate: Trace actual row states and rollback scope.
A payment-provider call times out while a seat hold approaches expiry. What state must exist before the call, and when may seats be released?
Reveal a model answer
Before calling the provider, I commit a PROCESSING hold state, a bounded processing deadline and one durable payment-attempt identity. A timeout leaves that attempt UNKNOWN; retries and reconciliation reuse it. Confirmation and expiry serialize on the hold and seats. The processing deadline may release inventory under the agreed policy even while the provider result is unknown, but a later successful charge must be compensated rather than taking those seats back.
Interviewer follow-up
What if the provider reports success after the processing deadline released the seats and another hold now owns them?
Reveal the follow-up answer
The verified success handler rechecks the old hold and every seat under locks. The ownership predicate fails, so it records the payment outcome and unique refund work without changing inventory. The new allocation survives. The external charge and local booking are separate transaction boundaries, so this compensation outcome is part of the product contract.
What the answer must demonstrate: Timeout is an uncertain external outcome.
Why does a FIFO waiting list not by itself guarantee fair booking?
Reveal a model answer
A newcomer could bypass the queue unless the hold endpoint consumes a current persisted admission grant. For the strict policy here, one unconsumed grant is active per contending show lane, issued to the oldest eligible ticket. The next ticket advances only after consumption, cancellation or expiry. Issuing several grants in order would not ensure redemption order, because a later customer’s network request might reach inventory first.
Interviewer follow-up
What if the first waiter wants six adjacent seats and the next wants one?
Reveal the follow-up answer
That is a product choice about head-of-line blocking. Strict FIFO may reduce utilization; allowing bypass changes fairness. I would define the policy explicitly rather than claim both unconditionally.
What the answer must demonstrate: Enforce queue policy where inventory is claimed.
The reservation service and waiting service restart together. What is recovered?
Reveal a model answer
I reconstruct active holds from their durable state and deadlines, and waiting order from persisted ticket sequence and session expiry. The current coordinator resumes grants for each show; the database rejects grants from a replaced coordinator. Browser notifications may replay, so clients read authoritative status and tolerate duplicate messages.
Interviewer follow-up
Can you reconstruct waiting order from current seat rows?
Reveal the follow-up answer
No. Seat state does not contain who arrived first or canceled. Waiting tickets need their own durable records; replicas of unrelated inventory do not preserve that information.
What the answer must demonstrate: Name the missing state, not just replicas.
Why shard by show instead of movie?
Reveal a model answer
Seat contention is scoped to one show, and its seats should share a transaction domain. Movie partitioning places every showing of a blockbuster on one owner unnecessarily. Show partitioning distributes different performances while preserving local seat-set transactions.
Interviewer follow-up
Does that make one sold-out opening show infinitely scalable?
Reveal the follow-up answer
No. The same scarce seats remain contended. I bound admission, cache browsing, and serialize or reject excess reservation work rather than hide the bottleneck behind a hash function.
What the answer must demonstrate: Distinguish distributed throughput from one-resource contention.
Hold H7 expires and releases seats 54–56. Hold H8 then acquires 56–57 before H7’s payment reports success. What exact condition prevents the late callback from stealing seat 56?
Reveal a model answer
The success transaction locks H7 and its immutable seat set, then requires H7 to remain PROCESSING before its processing deadline and every seat to remain allocated to H7. Because H7 is released and seat 56 belongs to H8, that predicate is false. The handler records the verified payment result and inserts unique compensation work without restoring H7 or changing H8’s seats.
Interviewer follow-up
Why is validating the hold before the provider call insufficient?
Reveal the follow-up answer
The network call occurs outside the database transaction and can outlast the deadline. Expiry or cancellation may change eligibility while it runs, so confirmation must recheck current hold state, deadline and all seat owners at the final local commit.
What the answer must demonstrate: State the atomic predicate and the losing outcome.
Two waiting-list coordinators believe they own the same show. Where must the system reject a stale coordinator’s admission grants?
Reveal a model answer
The show authority stores one active coordinator epoch. Every grant-creation transaction checks that epoch while advancing the durable waiting sequence. A replacement obtains a newer epoch through an atomic update; an old coordinator cannot create a valid grant with the previous epoch. Hold creation then validates and consumes the persisted grant in the same inventory transaction. A signed token alone does not establish that it is current or unused.
Interviewer follow-up
Does that solve two writable database primaries?
Reveal the follow-up answer
No. The storage failover layer must prevent split-brain writes. Application epochs help reject stale coordinators only when both reach one authoritative durable state.
What the answer must demonstrate: Identify the enforcing store and the remaining failure boundary.
Blank-page exercise · 45 minutes
Build the answer yourself
Design exact-seat booking with five-minute holds and a bounded payment-processing grace period. For one show, interleave requests for seats 54–56 and 56–57, then handle an unknown payment outcome at expiry while another customer waits. Prove no oversell and all-or-nothing allocation.
- Model show-seat uniqueness and all-or-nothing orders.
- Distinguish advisory browsing, durable holds, and row locks.
- Trace the two overlapping transactions.
- Resolve payment/expiry using explicit versions and deadlines.
- Recover fair waiting and unknown payment state.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a ticket-booking serviceWhat persists during payment?Recall first, then reveal
A durable expiring hold; database row locks have already been released.
Lock briefly, hold durably.
Return to lessonDesign a ticket-booking serviceWhat does all-or-nothing mean?Recall first, then reveal
The order receives every requested seat or none. A conflict rolls back the complete requested set rather than leaving an unnoticed partial reservation.
One order, one atomic seat set.
Return to lessonDesign a ticket-booking serviceWhat does a payment timeout mean?Recall first, then reveal
The result is unknown; reconcile the same attempt before taking a new financial action.
Unknown is not declined.
Return to lessonFinal revision
Summary and interview notes
A seat map suggests availability; only a committed hold reserves the whole requested seat set. Short transactions control confirmation, cancellation and expiry. If payment succeeds after the seats were released, record refund work without taking seats from another customer.
Remember these points
- A durable hold lasts minutes, while its database locks exist only for short state transitions.
- Claim the request identity before allocating seats so concurrent identical requests recover one result.
- Checkout records one payment attempt and bounded processing deadline before contacting the provider.
- Confirmation and expiry use the same lock order, fresh deadline time and current ownership of every requested seat.
- Strict waiting priority requires enforcement at grant redemption; FIFO grant issuance alone does not preserve allocation order.
Interview tips
- Interleave overlapping requests for seats 54–56 and 56–57 and show the losing transaction leaves no partial hold.
- Pause the provider response past expiry, allocate a released seat elsewhere, then show why late success creates only refund work.
- Separate map QPS, admitted hold capacity and strict waiting-lane throughput.
Important qualifications
- Webhook authenticity is separate from matching the stored provider attempt, amount, currency and successful outcome.
- Provider idempotency retention is finite and does not replace durable business attempt history.
- A single show remains a contention domain; sharding different shows or adding replicas does not remove that limit.
Technical references
- PostgreSQL explicit lockingExplains row locks, conflicts, and deadlock considerations for short reservation transactions.
- Stripe idempotent requestsProvider-specific retry contract; business retention/reconciliation must still be designed.
- Stripe webhook event handlingDocuments signature verification, duplicate events and non-guaranteed delivery ordering; application reconciliation remains necessary.
- PostgreSQL INSERT and conflict handlingSupports unique request claims and conflict-aware insert behavior; the booking transaction still defines all-seat atomicity.
- PostgreSQL current date/time functionsExplains why deadline validation after a lock wait cannot rely on transaction-start now().
Practice marks stay in this browser.