100+ interview questions with worked answers and follow-ups. Select a topic, try your answer, then compare the reasoning. Checking a box records your own assessment; it does not grade your answer.
What is system design, and how would you begin “design file sharing”?
Reveal a model answer
“System design defines the components, stored data, interfaces and interactions needed to meet requirements. For file sharing I first ask who uploads, who downloads, size limits, and whether links are public, private, expiring or revocable. Then I agree on volume and what success means, and trace one upload before adding capacity.”
Interviewer follow-up
Should you ask twenty questions before drawing?
Reveal the follow-up answer
No. Resolve the few ambiguities that change the first design, state reasonable assumptions for the rest, and validate them while drawing. Excessive questioning can prevent you from demonstrating a working solution.
What the answer must demonstrate: Connect each clarification to an architectural consequence.
What is a correctness invariant? Give one for an upload-and-download API.
Reveal a model answer
“An invariant is a condition the system must preserve. Here, an incomplete upload must never become downloadable. I represent upload state explicitly and allow downloads only after completion is verified. The download handler can enforce this rule by checking the stored upload state before serving the file.”
Interviewer follow-up
Is “the service is fast” an invariant?
Reveal the follow-up answer
“Fast” needs a performance target and measurement window. The upload rule applies to every relevant request: verify the unchanging object version, then atomically mark it ready only if its metadata still permits that change. A completion arriving after deletion must leave the file deleted. The file store and metadata database do not share one transaction.
What the answer must demonstrate: Give an enforceable rule, not an adjective.
“It makes the complete request and stored state understandable. I can show which changes must succeed together, verify the rules that keep the data correct, and measure capacity. I split or replicate components when a workload, reliability requirement, or ownership boundary creates a reason, rather than assuming that a distributed diagram is inherently better.”
Interviewer follow-up
What if the interviewer immediately requires global scale?
Reveal the follow-up answer
I still explain the logical operation, then show regional routing, data ownership, and replication. Starting from a clear operation does not require deploying only one machine.
What the answer must demonstrate: Logical clarity should survive changes in physical scale.
A file service averages 2.3 downloads/s but may peak at 1,000/s. Is average QPS enough to choose one server?
Reveal a model answer
“Not from that average alone. I need peak request rate, average and large-file sizes, connection duration, and a per-server load test at the target latency. Traffic can be concentrated into short bursts. I would state the peak assumption and size for it, including a server failure.”
Interviewer follow-up
If metadata QPS is modest but downloads total 12 TB/month, what motivates separate object storage?
Reveal the follow-up answer
Byte delivery dominates metadata traffic: the example has 12 TB/month of downloads. Separating bulk bytes gives an independent delivery path even if metadata QPS is modest.
What the answer must demonstrate: Do not equate average QPS with capacity.
Why define a data model before naming a database product?
Reveal a model answer
“The model tells me what must be stored together and which queries must be efficient. For file sharing I need ownership, upload state, expiry and the public-token mapping checked by identifier. A transactional metadata database can enforce those relationships; I evaluate products after deciding durability, throughput and failure requirements.”
Interviewer follow-up
When might that choice change?
Reveal the follow-up answer
Measured limits or global write requirements might justify partitioning or distributed storage. I would describe the new operational and consistency costs alongside the change.
What the answer must demonstrate: Explain access patterns and constraints.
An upload completion commits but its response is lost. How should a retry behave?
Reveal a model answer
“A timeout means the client does not know the outcome. I keep a stable upload identifier and make the completion operation inspect its existing state. Retrying completion for an already-ready upload returns the same worksheet. I recover the existing outcome before creating a new upload.”
Interviewer follow-up
How would you prove the fix works?
Reveal the follow-up answer
Simulate losing the response after the state transition commits, retry with the same identifier, and verify that exactly one logical worksheet is ready.
What the answer must demonstrate: Explain what happens if the server saves the result but the response is lost.
How do you answer “Why a CDN?” without a buzzword list?
Reveal a model answer
“Repeated downloads request identical bytes. A CDN can reduce origin traffic and serve a nearby copy. I would use versioned public objects where possible. If a link is private or revocable, I must define the authorization and cache lifetime so an old edge copy cannot bypass the promised access policy.”
No. Metadata still owns upload state, ownership, and permission decisions. Depending on the access scheme, the edge may enforce a limited authorization token or request validation.
What the answer must demonstrate: Name the benefit and the access-policy cost.
“I would recap the agreed user actions, trace the main path briefly, and state the key choices: durable upload states, independent byte delivery, and retryable completion. Then I would identify the first measured scaling limit and one remaining risk, such as revocation latency, with how I would test it.”
Interviewer follow-up
What if the design is unfinished?
Reveal the follow-up answer
Identify the part of the design you have not resolved and explain the next concrete decision you would make. A coherent partial design with clear invariants is better evidence of understanding than pretending every problem is solved.
What the answer must demonstrate: Summarize decisions and limits rather than reciting components.
“No. I can run multiple instances of the same application when its durable state is shared appropriately. I would extract a capability when independent capacity, releases or ownership justify the extra coordination. In checkout, keeping orders and stock reservations together preserves a useful local transaction; catalog search can scale separately as a derived view.”
Interviewer follow-up
What changes if inventory becomes a separate service?
Reveal the follow-up answer
Order creation and inventory reservation no longer share the original database transaction. The workflow must record progress, identify retries and recover failed or uncertain steps. A reservation needs explicit expiry and confirmation rules. An API call alone does not make the two services commit together.
What the answer must demonstrate: Distinguish server count, deployment boundaries and transaction boundaries.
What is an API, and what happens in a GET /orders/O17 request?
Reveal a model answer
An API is a contract between programs for an operation and its inputs, results, and failures. For GET /orders/O17, the client resolves the service name and establishes or reuses a protected connection. The edge routes the request; the order service validates the credential, derives user U9, checks U9’s permission for O17, and returns an authorized representation.
DNS supplies endpoint-discovery information such as an IP address. It does not fetch order data, authenticate the caller, or perform the database operation.
What the answer must demonstrate: Define the contract before tracing the complete request path.
“The token identifies or authorizes the caller, but a plaintext network could expose or alter it. TLS protects the communication and authenticates the server endpoint. I still validate the token and resource permission inside the service.”
Interviewer follow-up
If TLS ends at the edge, is the backend hop protected?
Reveal the follow-up answer
Only if I explicitly protect that hop too. Edge termination and backend encryption are separate connections.
What the answer must demonstrate: Separate transport protection from access checks.
Does HTTP/3 eliminate head-of-line blocking everywhere?
Reveal a model answer
“No. QUIC avoids TCP’s cross-stream loss-delivery blockage, but each stream still has ordering requirements and the application, queues, or shared resources can block progress. I choose it for actual transport needs, not as a blanket latency guarantee.”
“Safe methods do not ask for a state-changing action. Idempotent methods have the same intended effect when repeated. A deletion can be idempotent while still changing state. I do not use a GET to trigger a purchase merely because it is easy to call.”
Interviewer follow-up
Can POST be retry-safe?
Reveal the follow-up answer
Yes, if the application binds a stable operation key to one request and its stored outcome. HTTP method choice alone does not implement that protocol.
What the answer must demonstrate: Explain intended effect, not identical response bytes.
The request was accepted for processing, not completed. For our recoverable API, I durably commit an operation record and outgoing intent before 202, then return an operation ID and status location. HTTP 202 alone does not establish that storage guarantee.
Interviewer follow-up
What if the worker later fails?
Reveal the follow-up answer
The durable operation state must expose failure or a retry/recovery state. A successful enqueue is not the same as successful business completion.
What the answer must demonstrate: Distinguish acceptance and completion.
“I choose from client compatibility, schema tooling, and streaming needs. A public browser-facing API may use resource-oriented HTTP/JSON; internal typed calls may use gRPC. Both still need deadlines, authorization, and a defined retry contract.”
“A stable position in the chosen order, such as the last creation timestamp plus a unique order ID. The service validates it, applies the same ordering, and caps page size. I also define whether new or deleted records can change later pages.”
Interviewer follow-up
Why isn’t timestamp alone always enough?
Reveal the follow-up answer
Two orders can share a timestamp, so add a unique order ID. For descending history, continue strictly below the last (createdAt, orderId) pair. A cursor does not freeze the rows between requests; reproducible exports need a retained snapshot or saved result set.
What the answer must demonstrate: Match the cursor to the index and contract.
How do you rename a required response field safely?
Reveal a model answer
“I cannot assume all clients update together. I might serve both fields during migration or introduce a versioned contract, measure adoption, and retire the old field under an explicit policy. I test mixed client/server versions.”
Interviewer follow-up
Are all added fields automatically safe?
Reveal the follow-up answer
Only if existing consumers tolerate unknown fields and the new field does not change required semantics. Strict deserializers or signature schemes can make even additive changes significant.
What the answer must demonstrate: Describe a mixed-version rollout.
What is capacity estimation? Estimate QPS for one million users making ten requests a day.
Reveal a model answer
“Capacity estimation converts a workload into rates and resource needs. Here one million users × ten requests is ten million requests/day. Dividing by 86,400 seconds gives about 116 requests/s average. I still need peak concentration, bytes per request, latency targets and failure headroom before choosing server capacity.”
“Yes. A batch worker may finish thousands of items a second while each item waits minutes in a queue. Throughput describes the completion rate; latency measures one item’s elapsed time. I would measure queue wait and processing time separately.”
How much storage do 200 GB/day of uploads need after a year?
Reveal a model answer
“Without deletion, 200 × 365 is 73,000 GB, or 73 TB in decimal units. That is logical originals. I would separately add derived images, indexes, copies, and backups, then apply the retention policy.”
What happens at 2,000 requests/s if average latency grows from 50 to 500 ms?
Reveal a model answer
“Assuming both measurements cover the same system in stable operation, the average number of requests in progress grows from about 100 to 1,000. That can exhaust memory or connection pools even without a traffic increase. I would inspect downstream latency and bound admitted work.”
“Requests are not stored objects. I estimate distinct hot keys and bytes per entry. If a million requests hit one record, that is one cache entry. I use observed reuse and eviction behavior to choose the working set, then account for replication and overhead.”
Interviewer follow-up
How can a 99% hit ratio be misleading?
Reveal the follow-up answer
The remaining 1% may be huge objects or expensive queries. Measure byte hits, expensive misses, and cold-cache behavior.
What the answer must demonstrate: Count distinct retained entries.
Three servers can just meet peak. Is that a resilient design?
Reveal a model answer
“Not if the requirement includes surviving a server failure at that peak. I calculate the remaining capacity after the failure and keep headroom for imbalance. If two survivors cannot meet the objective, I add capacity, reduce admitted work, or agree on degraded behavior.”
Interviewer follow-up
Can autoscaling replace all spare capacity?
Reveal the follow-up answer
Autoscaling has detection and startup delay, and dependencies may scale more slowly. A sudden failure needs capacity or load shedding during that interval.
What the answer must demonstrate: Calculate surviving capacity.
An API runs five database queries. Which QPS matters?
Reveal a model answer
“Count both. At 2,315 API requests/s and five queries per request, the database receives about 11,575 operations/s before retries or cache effects. I would check whether each query is needed, indexed and independent of the others.”
Interviewer follow-up
What if all five run in parallel?
Reveal the follow-up answer
Parallel execution can lower one request’s latency, but it still creates roughly five operations of load. Latency and total work are different.
What the answer must demonstrate: Explain amplification rather than hiding it.
A photo service serves 2,315 peak views/s at 100 KB each. Which measurements would change the storage or delivery design?
Reveal a model answer
“The stated peak is 2,315 × 100 KB = 231.5 MB/s, about 1.85 Gb/s before overhead. I would measure repeated-key reuse and permission constraints to evaluate a CDN, and measure metadata and CPU costs separately to decide where scaling helps. Peak QPS alone cannot determine daily delivered bytes or metadata growth; those require daily volume and stored bytes per upload.”
Interviewer follow-up
What observation could invalidate your CDN assumption?
Reveal the follow-up answer
If most images are viewed only once or authorization prevents useful sharing, cache reuse may be low. I would measure the access distribution and policy constraints.
What the answer must demonstrate: Use a number to justify a decision.
“Availability asks whether an eligible checkout operation can complete under its success definition. Reliability asks whether the service performs its specified function correctly over time and under promised conditions. A reachable system that double-charges an order is incorrect; that purchase must also count as unsuccessful in an end-to-end availability measure. I define the outcome and measurement window rather than treating reachability as either guarantee.”
Interviewer follow-up
Can you preserve correctness while losing availability?
Reveal the follow-up answer
Yes. Refusing a purchase when the service cannot confirm which node may update inventory avoids accepting an order it cannot safely reserve stock for, but the customer still cannot complete the operation.
What the answer must demonstrate: Use the same example for both qualities.
“If the database fits on one larger instance and measured CPU, memory, or I/O is the bottleneck, vertical scaling can buy capacity with a smaller operational change. I would also keep redundancy and test the new capacity. I shard when independent data needs to exceed that practical limit.”
Why does doubling application servers not double checkout throughput?
Reveal a model answer
“They may still share the same database, lock, or downstream service. I trace a purchase and measure where time and work accumulate. Adding application capacity helps only the work those instances own; the shared inventory writer may remain the limiting resource.”
Interviewer follow-up
What if browsing scales but checkout does not?
Reveal the follow-up answer
That is plausible because browsing can distribute read work while checkout changes shared inventory. I would size and design those operations separately.
What the answer must demonstrate: Find the shared bottleneck.
“First I would define the measure. Over a 30-day time-based window, 0.1% is 43.2 minutes. Over a million eligible requests, it is 1,000 unsuccessful attempts. These budgets are not interchangeable when traffic changes through the day.”
Interviewer follow-up
Can a fast error count as successful?
Reveal the follow-up answer
Only if it is a valid business response under the defined metric, not because the network responded quickly. An infrastructure refusal of a valid purchase is an unavailable outcome.
What the answer must demonstrate: Define eligible and successful requests.
“The customer experiences the outage before the operator starts repairing. If detection takes one minute and verified failover takes six more, checkout is unavailable for seven. I improve both detection and repair and practise the complete sequence.”
Interviewer follow-up
Would aggressive health checks always help?
Reveal the follow-up answer
No. Noisy checks can eject or restart live capacity during a shared dependency incident. Stop sending work to a known dead instance, but use bounded admission and careful failure thresholds to prevent overload from cascading through the survivors.
What the answer must demonstrate: Measure end-to-end recovery.
“No. I need to specify when a write is acknowledged, whether the second copy is durable, and which failures it survives. Copies in the same failure domain may disappear together, and a bad deletion can replicate to both. I also need backups and tested recovery.”
Interviewer follow-up
What is a failure domain?
Reveal the follow-up answer
A set of resources that can fail together because they share a dependency, such as power, a rack, a zone, or an administrative change.
What the answer must demonstrate: Name the failure being tolerated.
“No. One message may contain a huge unused payload, while several small messages may run in parallel. I compare bytes, round trips, CPU, and end-to-end latency for the same user operation. Reducing repeated calls can help, but the workload decides.”
Interviewer follow-up
What changes across regions?
Reveal the follow-up answer
Each sequential round trip can cost substantially more time because of distance. I would reduce cross-region dependencies on the critical path and measure the actual network.
What the answer must demonstrate: Count bytes and sequential waits, not just arrows.
What makes a system manageable in an interview answer?
Reveal a model answer
“I show how an operator diagnoses one failed order using a trace identifier and durable states, how alerts reflect failed purchases, and how a rollout can be stopped or reversed. I include schema compatibility and verify recovery rather than ending the design at deployment.”
Interviewer follow-up
Which metric would you alert on first?
Reveal the follow-up answer
The user-facing purchase-success or latency objective, supported by component metrics to locate the cause. A low-level CPU signal alone does not establish customer impact.
What the answer must demonstrate: Explain a concrete operator action.
What are a data model, access pattern, invariant, and transaction? How do they guide database choice?
Reveal a model answer
A data model describes the representation: tables, documents, key-value pairs, or graph relationships. An access pattern is a specific query or update, such as recent orders for customer U7. An invariant is a rule that must remain true, such as stock never becoming negative. A transaction treats operations as one logical unit whose changes commit or roll back together. ACID names atomicity, ACID consistency, isolation and durability; the engine and its settings determine the exact guarantees.
For an order service, write down order-by-ID, customer history, and conditional stock allocation. If reducing stock from 5 to 3 must commit with creating a $24 order, a relational database with suitable indexes and a local transaction is a straightforward starting point. Then test expected volume, hot-item contention, and the actual engine's features. SQL and NoSQL labels alone do not determine scale or transaction support.
Interviewer follow-up
Would a billion rows automatically change your choice?
Reveal the follow-up answer
“No. Row size, access locality, indexes, request rate, and partitioning capabilities determine the bottleneck. I would identify the limiting operation first.”
What the answer must demonstrate: Size alone does not describe a workload.
An order contains its item lines, while product inventory is shared across many orders. Would storing each order as one document make the whole purchase atomic?
Reveal a model answer
“Embedding O81’s lines makes the order read convenient, but MUG9 stock is shared by many orders. Copying available quantity into each order creates competing truths. I would keep stock in one authoritative inventory system and use a supported transaction, or an explicit reservation workflow, to coordinate stock allocation with the order.”
How does wide-column differ from analytical columnar storage?
Reveal a model answer
“A wide-column model can place U7’s orders in one partition and order them by time for a known serving query. Analytical columnar storage supports scans of selected attributes across many records. Similar names do not make their access shapes or guarantees interchangeable.”
Interviewer follow-up
Where would a monthly aggregate report run?
Reveal the follow-up answer
“I would consider a derived analytical path if scans disrupt purchases, then define its lag and reconciliation. A serving database and report workload need not share one bottleneck.”
What the answer must demonstrate: Avoid treating column-related names as one category.
A purchase must create order O81 for two $12 items and reduce stock from 5 to 3. Explain ACID for that transaction.
Reveal a model answer
“Atomicity makes stock allocation and order insertion succeed together or have neither change take effect. Correct logic preserves nonnegative stock. Isolation governs concurrent buyers. Durability defines which failures committed O81 survives. I would show the transaction and its settings because saying ‘ACID database’ does not prove the application rule.”
No. ACID consistency preserves database and application rules, such as nonnegative stock. CAP consistency means linearizability: after a write completes, a later read must see it or a newer write. A store can serve fresh values while bad transaction logic breaks a business rule.
What the answer must demonstrate: Name the rule and distinguish the two meanings.
Two concurrent purchases each request two units when stock is two. What prevents overselling?
Reveal a model answer
“I put UPDATE Inventory SET available = available - 2 WHERE sku = the_requested_sku AND available >= 2 in the same transaction as the order insertion, and require one affected row before continuing. In PostgreSQL Read Committed, the second updater waits and rechecks the predicate. If the first commits stock 2 → 0, the second affects zero rows and rolls back instead of creating an order. A stock CHECK constraint is useful defense, but I still need the transaction and affected-row check.”
Interviewer follow-up
What if the rule spans several products?
Reveal the follow-up answer
“I need a transaction strategy protecting the whole rule or a deliberate reservation workflow. One row’s condition cannot enforce an unstated cross-row invariant.”
What the answer must demonstrate: A fresh read is not an atomic allocation.
“Old and new consumers must agree on quantity, currency, and record versions. Permitting multiple shapes does not tell the application how to interpret them. I would validate required fields and stage compatible readers and writers so a storage change does not silently change meaning.”
Interviewer follow-up
Must a relational schema alteration require downtime?
Reveal the follow-up answer
“Not universally. The exact operation and engine determine locks and rewrite costs; many changes can be staged compatibly.”
What the answer must demonstrate: Flexibility does not eliminate migration work.
Order O81 commits but the response is lost. How should the application recover the outcome?
Reveal a model answer
“The retry carries the same customer-scoped purchase key and request. I claim that unique key when inserting the uncommitted order, before allocating stock. If the key conflicts, I roll back the attempt, then use a fresh transaction to read and validate the original order’s request hash. This returns the original success even if it exhausted the remaining stock. A new purchase ID or a stock check performed before resolving the duplicate would give the wrong retry behavior.”
Interviewer follow-up
What if the retry changes the quantity?
Reveal the follow-up answer
“I reject reuse of the same identifier for different request data, or apply an explicit documented policy. It cannot silently mean another purchase.”
What the answer must demonstrate: Unknown commit is different from known rollback.
What changes when inventory becomes an independent service?
Reveal a model answer
“The stock and order updates no longer share the original local transaction. I must choose a distributed transaction or durable reservation workflow with explicit intermediate and compensation states. Moving tables across owners without revisiting that boundary loses the guarantee my first design depended on.”
“It is a maintained search structure that helps locate records without checking every row. An author catalog points to books by an author. In a database, the index stores searchable keys and enough information to find or return matching data.”
Interviewer follow-up
Why not create one for every field?
Reveal the follow-up answer
Every maintained index consumes space and adds work to relevant inserts, updates, and deletes. I choose indexes from actual queries and constraints.
What the answer must demonstrate: Explain the read/write tradeoff.
How does an index on (author, title, id) answer author = Le Guin ordered by title?
Reveal a model answer
“It seeks to the first Le Guin entry and scans that contiguous author range in title order. It fetches the matching book rows only if required fields or visibility checks need them, then stops at the range end or limit. The benefit is avoiding unrelated authors, not assuming every query can be served entirely from the index.”
Interviewer follow-up
Would it help a query on title alone equally well?
Reveal the follow-up answer
Not necessarily: author is the leading ordering. The engine may use another access method, but a title-leading index directly matches that different query.
What the answer must demonstrate: Walk the keys rather than naming the structure.
Which index fits customer history sorted newest first?
Reveal a model answer
“I start with customerId, then createdAt descending, then orderId descending for ties. Equality on customer narrows the range and the remaining order supports the requested slice. I would include returned columns only if reducing row lookups justifies a larger index.”
Interviewer follow-up
Why include orderId when timestamps exist?
Reveal the follow-up answer
Two orders can share a timestamp. A unique tie breaker creates a deterministic order and a complete pagination cursor.
What the answer must demonstrate: Explain equality, ordering, and tie breaking.
When inserting a new book row with ID 15, what additional work do maintained indexes require?
Reveal a model answer
“The table gets a row and each maintained index gets a corresponding entry. The storage engine also performs its logging and any page maintenance required. Extra indexes therefore increase write amplification, memory pressure, and storage even if this insert is only one business operation.”
Interviewer follow-up
What about an update to an indexed title?
Reveal the follow-up answer
The author/title index must reflect the new key. The exact update mechanism depends on the engine, but later queries must find the correct title for the row version they are allowed to read.
What the answer must demonstrate: Account for all maintained structures.
“It contains the fields needed to answer a query, potentially avoiding separate row fetches. For order history I might include total with the ordering keys. Whether an index-only scan is actually possible also depends on the engine’s visibility rules and query plan.”
Interviewer follow-up
What is the cost of including total?
Reveal the follow-up answer
More index bytes and maintenance when total changes. I measure whether saved reads justify that cost.
What the answer must demonstrate: Do not promise every covered query avoids all table access.
“The database may still walk past the earlier matching entries before returning the requested page. A keyset cursor lets the next query seek after the last seen ordering tuple. I use a stable tie breaker and define how concurrent inserts affect the browsing session.”
Interviewer follow-up
Does a cursor guarantee an unchanged snapshot?
Reveal the follow-up answer
No. It identifies a position. A snapshot across pages requires an additional consistency/version mechanism if the product needs it.
What the answer must demonstrate: Separate ordering and snapshot consistency.
“A query matching 90% of a table may do more work through index-to-row lookups than through a sequential scan; a query matching 100 rows in a million has a different cost. I inspect estimated versus actual rows, buffers, filtering and sort work for the exact query. Small tables, stale statistics and data skew can change the plan.”
Interviewer follow-up
Would a low-cardinality boolean index always be useless?
Reveal the follow-up answer
No. It can help selective partial queries or specific engine strategies. The useful question is how much work it avoids for this query and distribution.
What the answer must demonstrate: Avoid absolute rules disconnected from data.
“A local index searches within its storage owner. The request still needs to identify the right shard, or query a distributed index or multiple owners. For customer history, customer-based routing and a customer/time local index work together.”
What is a storage engine, and how is it different from a data model?
Reveal a model answer
The data model describes records and access semantics, such as messages keyed by room and sequence. The engine organizes their bytes and indexes and performs updates and recovery. B-trees and LSM trees are engine techniques; relational tables and documents are logical models. Choosing SQL does not by itself select a B-tree or define its disk cost.
Yes. A document representation does not force one physical engine. I compare the actual implementation’s transactions, indexes, durability, and maintenance work for the required queries.
What the answer must demonstrate: Distinguish the logical interface from physical organization.
A B-tree has separators 20 and 50; its middle leaf contains 21, 35, 42, 49. Explain lookup for key 42.
Reveal a model answer
“The root separators guide me to the relevant leaf range, where I find 42’s index entry. Depending on the layout, that entry contains the needed data or points to a separate row. Cached pages can avoid disk reads.”
Interviewer follow-up
Why not say it always takes three I/Os?
Reveal the follow-up answer
“The tree’s height, cached pages, record layout and overflow data all affect physical work.”
What the answer must demonstrate: Distinguish logical search steps from physical I/O.
An update is acknowledged before its changed data page reaches disk. Under what WAL policy can it survive a process crash?
Reveal a model answer
It can survive when the required recovery records, including the commit decision, were made durable before acknowledgment and recovery correctly replays them. Log-before-data ordering alone does not prove commit-before-ack durability. I must verify the configured synchronization policy and failure model.
Interviewer follow-up
What if the disk is destroyed?
Reveal the follow-up answer
“Then a local WAL alone is insufficient; the replica or backupdurability policy determines what survives.”
What the answer must demonstrate: Name the acknowledgment boundary and failure model.
“The old sorted file cannot be changed. An edit first enters a newer memory table and later another file. Reads use the engine’s sequence and snapshot rules to choose the right version. Compaction removes old versions once they are no longer needed.”
Interviewer follow-up
Can a read return the first copy it finds?
Reveal the follow-up answer
“Only if the search protocol proves it is the correct visible version. Arbitrary file traversal is not sufficient.”
What the answer must demonstrate: Explain version visibility, not just file count.
An LSM contains a tombstone for key 8 and older files may contain key 8’s value. When may the tombstone be removed?
Reveal a model answer
“Only when the engine can prove older values cannot reappear for supported reads and no required snapshot needs that history. Removing the marker merely because it is old can expose an older stored copy.”
“With a measurement that includes all the relevant local writes, it implies roughly 800 MB/s of device writes. I would also budget compaction reads, CPU, replication and headroom, and verify the figure under a steady workload.”
Interviewer follow-up
Can you compare two quoted amplification numbers directly?
Reveal the follow-up answer
“Only if their numerator, denominator, workload and inclusion of logs or replication match.”
What the answer must demonstrate: Define the measurement before multiplying it.
Why does an ordered engine not automatically give fast room history?
Reveal a model answer
“The logical key and partitioning still matter. If each full message key is independently hashed to a different shard, a room query fans out. Keeping room and sequence together gives locality but may create a hot room partition.”
Interviewer follow-up
What does bucketing change?
Reveal the follow-up answer
“It bounds one partition’s size or traffic, while making history retrieval merge results from several buckets.”
What the answer must demonstrate: Connect query shape to both ordering and partitioning.
“I would load representative data, sustain ingestion until compaction reaches normal behavior, and measure tail latency for latest-fifty reads, edits, deletions and recovery. An empty database’s short insert burst hides the deferred maintenance cost.”
Interviewer follow-up
What failure would make you reconsider the choice?
Reveal the follow-up answer
“A persistent backlog of compaction work or range-read latency percentiles exceeding the target would prompt layout and resource changes, or a simpler engine better suited to the actual workload.”
What the answer must demonstrate: Evaluate steady-state operation, not only peak foreground throughput.
“It chooses a healthy backend for incoming service traffic. The client uses one service address; the balancer can route its request to A or B. It distributes work and helps route around detected failures, but shared data and correct write ownership still need their own design.”
Interviewer follow-up
Does it automatically make the application stateless?
Reveal the follow-up answer
No. If the cart exists only in A’s memory, switching to B can lose it. The application must place essential state where another instance can recover or access it.
What the answer must demonstrate: Explain routing separately from state.
Where do six equal requests go across A, B, and C?
Reveal a model answer
“Under simple round robin: A, B, C, A, B, C. If A has twice the capacity, I can use weights giving A roughly half the traffic. These choices assume requests are similar enough that request count represents work.”
Interviewer follow-up
What breaks that assumption?
Reveal the follow-up answer
A large export can use far more CPU or time than a small read. Equal request counts may create uneven load. I would separate pools or consider a work-sensitive signal.
What the answer must demonstrate: Demonstrate a schedule before discussing limitations.
When should a service use L7 routing instead of L4 balancing?
Reveal a model answer
“I need to route by HTTP path, such as /images versus /checkout. An L7 balancer understands those fields, usually after TLS termination. I would also protect the backend connection. An L4 connection balancer is sufficient when I only need transport-level distribution.”
Interviewer follow-up
Can one HTTP/2 connection represent many requests?
Reveal the follow-up answer
Yes. Multiplexing means connection counts are not request counts, so a connection-level policy can still produce uneven application work.
What the answer must demonstrate: Describe what information the layer can inspect.
A has 10 connections, B 2, C 5. Who gets the next one?
Reveal a model answer
“Least connections chooses B if the servers and connection costs are comparable. I would not assume B is least busy if those two connections each contain many expensive streams. I validate the signal against CPU, queueing, and latency.”
Interviewer follow-up
Why not always use least response time?
Reveal the follow-up answer
It relies on measurements that may lag or react poorly to small samples. A recently idle slow node may look deceptively good, and routing can oscillate.
What the answer must demonstrate: Qualify the unit of work.
“They reduce movement while the chosen server works, but they do not preserve a cart when it fails. I store the authoritative cart durably and use stickiness only if locality improves performance. Then another backend can continue the session.”
Interviewer follow-up
Why is IP hashing risky behind a school network?
Reveal the follow-up answer
Many users can share one public NAT address and hash to the same backend. A client IP is not a unique user identifier.
What the answer must demonstrate: Separate locality and durability.
A backend B crashes before its next health probe. What happens until the balancer removes it?
Reveal a model answer
“Some requests may still be sent to B until failure detection crosses its threshold. I use bounded timeouts and safe retries. After removal, A and C must have capacity for the redirected work; otherwise detection can turn one crash into a broader overload.”
Interviewer follow-up
Should readiness check every downstream system?
Reveal the follow-up answer
Only dependencies necessary for that pool’s promised work. Active probes exercise a selected path; passive checks observe actual failures. Checking an optional shared service can eject the entire pool unnecessarily, while a shallow process probe can miss failed critical operations.
What the answer must demonstrate: Acknowledge detection delay and correlated failures.
“I stop new assignments, signal clients to reconnect where the protocol allows, and enforce a drain deadline. Message state lives in durable storage so reconnecting to another instance can resume from a cursor. I cannot assume a balancer transfers the old socket’s process memory.”
Interviewer follow-up
What if a client never disconnects?
Reveal the follow-up answer
The deadline eventually closes it. The application protocol must make reconnection and replay a supported path.
What the answer must demonstrate: Explain the long-lived session explicitly.
Does adding a load balancer eliminate all single points of failure?
Reveal a model answer
“No. The balancer, discovery and shared database are separate dependencies. I use redundant balancers with supported traffic failover and enough surviving capacity. If the failed balancer terminated a client connection, that connection may still break: the client reconnects and retries safely. Routing a new request is different from preserving an old connection.”
Resolvers and clients may keep cached answers until their expiration behavior permits a refresh. I would avoid promising universal instant traffic movement.
What the answer must demonstrate: Trace the full failure path.
What is caching? Use product P7 at $20/version 8 to explain the first miss and a subsequent hit.
Reveal a model answer
“Caching keeps a reusable copy to avoid repeating a more expensive operation. Request R1 for product:P7 misses, so the application loads $20/version 8 from the database and stores a copy. The next permitted request R2 hits that copy. The database remains authoritative; the hit is usable only under the page’s freshness and access policy.”
Interviewer follow-up
Does the cache update itself when the database changes?
Reveal the follow-up answer
Not generally. We need a write/update/invalidation mechanism or let an expiry trigger a later refresh.
What the answer must demonstrate: Name the source of truth.
When would you choose a local cache rather than a shared one?
Reveal a model answer
“A local memory cache is fast and avoids a network dependency; a local disk cache can hold larger reusable objects. But copies differ across application instances and vanish or become unavailable with the host. A shared cache simplifies sharing at the cost of a network call and another service to operate.”
Interviewer follow-up
What happens under random backend routing?
Reveal the follow-up answer
A repeat caller may land on an instance whose local cache has never seen the key. Hit rate depends on placement and request distribution.
What the answer must demonstrate: Explain per-instance copies.
Does a 30-second TTL guarantee every read is less than 30 seconds stale?
Reveal a model answer
“Only under specified fill, age, and refresh rules. If a delayed reader fills an already old value with a new 30-second timer, its data age may exceed that bound. I would carry version or source timestamps when the age limit matters and define which moment starts the TTL.”
Interviewer follow-up
What if permissions must change immediately?
Reveal the follow-up answer
I need current authorization or a revocation mechanism that enforces that promise. A general long-lived content cache cannot supply immediate revocation by itself.
What the answer must demonstrate: Distinguish cache residency age and data age.
Why can delete-after-write still return the old price?
Reveal a model answer
“Reader R can fetch version 8 before writer W commits version 9, then refill after W deletes the cache. The delete happened, but the late reader resurrected the old copy. I show that timeline and choose either bounded stale display or a stronger version-aware update protocol.”
Interviewer follow-up
How does a version floor help?
Reveal the follow-up answer
The cache records that P7 must be version 9 or newer and checks that rule atomically before accepting a refill. It then rejects version 8. Keep the rule while old refills may arrive; if it is lost, reject attempts from the old cache generation and rebuild safely. Stale reads remain possible between the database commit and installing the rule unless those steps are coordinated more strongly.
What the answer must demonstrate: Locate the late refill, then the atomic check.
Why not acknowledge orders from a write-back cache?
Reveal a model answer
“If the cache acknowledges before durable persistence and then loses the entry, the customer can lose an order already reported as saved. I would need a replicated durable log and a tested recovery protocol, or acknowledge only after the required durable commit.”
Interviewer follow-up
Can write-back ever be reasonable?
Reveal the follow-up answer
Yes, for a workload whose loss model permits it or a cache system designed to provide the required durability. The name of the pattern alone does not prove safety.
What the answer must demonstrate: Tie acknowledgment to a loss model.
A two-entry cache receives insert A, insert B, read A, insert C. What do FIFO and LRU evict?
Reveal a model answer
“With two entries and eviction from existing entries, FIFO evicts A because it was inserted first. LRU evicts B because A was accessed more recently. This demonstrates that insertion order and access order are different.”
A long scan of one-use entries can evict frequently useful keys. Admission policy or frequency-aware policies may help, depending on the access distribution.
What the answer must demonstrate: Replay the actual ordering.
Ten thousand readers miss P7 at once. What do you do?
Reveal a model answer
“I allow one refresh for P7 and coalesce the other requests behind it, with bounded waiting. If the product permits it, I serve a stale copy during refresh. I also limit database fallback globally so many different missing keys cannot overwhelm it.”
How would you add a CDN to an existing image service?
Reveal a model answer
“I keep static objects behind a stable static hostname and point delivery through the CDN. I set origin access, TLS, cache headers, and versioned object paths. On a miss the edge fetches the origin; on a permitted hit it returns its copy. Private objects need a separate authorization-compatible plan.”
Interviewer follow-up
Can you cache a personalized product response under only product ID?
Reveal the follow-up answer
No, if tenant, currency or authorization changes the response. Represent safe variation in the key, or exclude the response from shared caching with an appropriate policy. HTTP private prevents shared storage; no-cache permits storage but requires validation. Neither substitutes for checking access.
What the answer must demonstrate: Explain both migration and key correctness.
What is the difference between a forward and reverse proxy?
Reveal a model answer
“A forward proxy represents clients reaching external destinations, such as employees using a company web gateway. A reverse proxy represents servers to incoming callers, such as shop.example forwarding to an internal order service. Both relay responses back; the names describe role, not one-way packet direction.”
Interviewer follow-up
Can one gateway product perform both?
Reveal the follow-up answer
A product can support several roles, but I still configure and explain the client trust and destination policy for each deployment.
What the answer must demonstrate: Say whose behalf the proxy acts on.
When a reverse proxy terminates HTTPS for an order API, which connection does TLS protect?
Reveal a model answer
“The browser’s TLS connection ends at the reverse proxy, which presents the shop certificate and can inspect the HTTP request. The proxy may then establish a separate protected backend connection. I would not assume browser-to-edge encryption automatically protects the entire path.”
The intermediary forwards encrypted traffic and cannot inspect the HTTP path or headers inside it. A forward proxy can similarly establish a CONNECT tunnel and then relay browser-to-origin TLS. In either case, destination policy and transport metadata remain separate from decrypted application content.
What the answer must demonstrate: Draw both connection segments.
Why can’t the backend trust any X-Forwarded-For value?
Reveal a model answer
“An external caller can send that ordinary header. In a one-edge deployment I replace untrusted claims at the edge with its observed address, and the backend accepts forwarding metadata only from that trusted edge. Multiple proxies require an explicit trusted-hop traversal rule. A forged localhost value must not grant internal access, and an IP address still does not establish user identity.”
Interviewer follow-up
Does the client IP establish the user identity?
Reveal the follow-up answer
No. Users can share addresses, and addresses can change. Authentication and resource authorization need separate credentials and checks.
What the answer must demonstrate: Separate network provenance and identity.
Can a reverse proxy share an authenticated order response using only its URL as the cache key?
Reveal a model answer
“No. A private order response must not become another user’s response. I choose an authorization-compatible cache policy, often avoiding shared caching for this path. Public versioned product images can use a different policy.”
Interviewer follow-up
Would including a user ID in the cache key be enough by itself?
Reveal the follow-up answer
It helps separate entries, but I still validate the identity and permissions, handle revocation, and ensure untrusted input cannot choose another user’s key.
What the answer must demonstrate: Protect the authorization decision as well as key separation.
“No. Open describes who is allowed to use it; anonymous describes which identifying information it tries to hide. A proxy can be open and still log users or forward identifying headers. Neither term alone establishes privacy or safety.”
Interviewer follow-up
What does transparent mean?
Reveal the follow-up answer
Clarify whether it means interception without client configuration, or historical HTTP forwarding of requests and responses without transformations beyond proxy authentication and identification.
What the answer must demonstrate: Treat role, access, and visibility as different dimensions.
The gateway times out on POST /checkout. Can it retry automatically?
Reveal a model answer
“Only if the checkout protocol makes repeating that logical request safe. The origin may already have committed the purchase while the response was delayed. A stable idempotency key and saved result let a retry recover the outcome; an arbitrary new POST may create a second purchase.”
The origin works but the public subpath fails. What do you inspect?
Reveal a model answer
“I inspect path stripping, relative links, redirects, query strings, and asset routes. If the origin redirects to a root-relative path, it may omit the public prefix. I verify the actual public URL rather than treating origin success as end-to-end proof.”
Interviewer follow-up
What is a concrete fix?
Reveal the follow-up answer
Rewrite relevant redirect locations at the proxy or publish links to canonical paths that preserve the public base, then test navigation and assets on both paths.
What the answer must demonstrate: Follow the visible URL through the proxy.
Which responsibilities would you keep out of a generic gateway?
Reveal a model answer
“I can centralize routing, TLS, request-size limits, and some authentication or quota checks. The order service must still check who may read or change an order, and the component committing a purchase must enforce rules such as not selling more stock than is available. Otherwise an alternate internal caller could bypass the only business check.”
Interviewer follow-up
Does adding two gateways remove all risk?
Reveal the follow-up answer
No. They may share one bad configuration or a saturated dependency. I also need safe rollout, monitoring, and surviving capacity.
What the answer must demonstrate: Explain responsibility and shared failure modes.
“A shard owns a different subset of records; a replica is another copy of the same records. A and B split customers, while A1 and A2 could be copies of shard A. I need separate rules for routing to an owner and for keeping that owner’s copies consistent.”
“The dominant query asks for one customer’s orders. Keeping those records together allows one routed query and local updates of related order data. I would verify the customer traffic distribution and identify global queries that this choice makes more expensive.”
Interviewer follow-up
Would order ID be equally good?
Reveal the follow-up answer
It can spread individual orders better, but listing a customer’s orders needs a secondary location/index path or fanout. The best key depends on the required queries.
What the answer must demonstrate: Connect the key to an actual query.
“Queries over adjacent keys can target a small set of contiguous ranges. It is useful when the range matches the query, such as a time slice. The risk is skew: always appending to the newest timestamp range can concentrate writes.”
Interviewer follow-up
Is every horizontal partition a range partition?
Reveal the follow-up answer
No. Horizontal means splitting records; hash, list, and other placement rules are alternative ways to do that.
What the answer must demonstrate: Explain the category and the method.
Hashing is uniform. Why is one shard still overloaded?
Reveal a model answer
“Uniform placement distributes keys, not necessarily requests. One customer may account for half the work, or one key may be exceptionally large. I inspect traffic and bytes by key, then consider splitting that workload, replicating reads, or allocating dedicated capacity.”
Interviewer follow-up
Can you split a customer without cost?
Reveal the follow-up answer
It can turn a formerly local order listing or transaction into cross-partition work. I explain that cost and preserve the required ordering or atomicity explicitly.
What the answer must demonstrate: Do not promise hashing eliminates hot keys.
What happens to a join between orders and products?
Reveal a model answer
“If they live on different owners, a local SQL join may no longer cover them. I can perform bounded application lookups, co-locate relevant data, or keep a suitable read copy. For receipts, recording product name and price at purchase time is often the correct historical data.”
Interviewer follow-up
Does denormalization mean every copy must stay current?
Reveal the follow-up answer
No. A historical purchase snapshot should remain historical; a current product description needs an update policy. Similarly, a local index proves uniqueness only in its own scope: a global email claim or order ID requires an explicit cross-shard constraint or single claim owner.
What the answer must demonstrate: Distinguish historical facts from current replicas.
“Keep B accepting writes while copying a consistent snapshot tied to log position L0. Apply later logged changes at C. To switch, stop B’s writes and make C apply through B’s final committed position. Then enable C under a new routing version and reject writes using B’s old version. Stale clients refresh their routes and retry the same operation. If I cannot prove B can no longer commit writes, I do not enable C.”
Interviewer follow-up
Why not switch the directory halfway through copying?
Reveal the follow-up answer
C may lack records or writes that arrived after the snapshot. The directory must not send writes to C until C has the required data and B can no longer accept conflicting writes.
What the answer must demonstrate: Separate data catch-up and ownership transfer.
“Clients can use a cached version only while the ownership protocol makes stale routes safe. Owners validate epochs and reject invalid writes. For metadata changes I need a durable authoritative directory; guessing a new owner can create conflicting histories.”
“Customer-based sharding does not localize a global time query. I can fan out bounded queries and merge results for modest needs, or stream order changes into an analytical store partitioned for reporting. I state the reporting freshness delay and avoid making every checkout wait for analytics.”
Interviewer follow-up
What if the report must be an exact cross-shard snapshot?
Reveal the follow-up answer
That needs a defined consistent snapshot or coordinated read protocol. Independently querying owners at different times does not automatically represent one instant.
What the answer must demonstrate: Name the cost of a query the key does not serve.
What is consistent hashing? Draw a ring and explain why adding a node moves fewer keys than changing a modulo divisor.
Reveal a model answer
Consistent hashing is a placement scheme that limits remapping when owners join or leave. Draw a ring numbered 0–99 with A at 20, B at 50, and C at 80. Hash a key and choose the first clockwise owner, wrapping at 99. Hash 35 belongs to B50; hash 90 wraps to A20.
Add D40: it takes only (20,40] from B, so hash 35 moves to D while hash 45 stays at B. Changing hash(key) mod 3 to mod 4 would change many unrelated assignments. In balanced equal-capacity placement, adding one to N owners moves about 1/(N+1) of keys on average; this particular D40 interval covers 20% of our toy ring. Virtual nodes improve balance, but data still needs migration or cache refill, and one hot key remains a separate problem.
Interviewer follow-up
What happens for hash 90?
Reveal the follow-up answer
“It wraps through 99 and 0 to A20. That wraparound interval is part of A’s ownership.”
What the answer must demonstrate: Demonstrate the rule with actual positions.
On a 0–99 ring with A20, B50, C80 and keys at 12, 35, 45, 65, 90, which keys move when D40 joins?
Reveal a model answer
“Only P35 moves in our five-key sample. D takes (20,40] from B; P45 is outside that interval and stays with B. A and C keep their existing intervals. I would show the interval, not claim that every key moves to a new server.”
Interviewer follow-up
Does moving the sample key at 35 imply exactly one quarter of all keys moved?
Reveal the follow-up answer
“No. In this fixed ring D40 receives (20,40], which is 20 of 100 positions. An expected 25% movement requires four balanced owners and suitable hash-distribution assumptions; a five-key sample need not match either fraction.”
What the answer must demonstrate: Keep a concrete trace distinct from a statistical estimate.
On a clockwise ring with A20, D40, B50, C80, which owner receives B50’s interval when B is removed?
Reveal a model answer
“B’s remaining interval (40,50] passes to C80, the next clockwise owner. P45 moves to C. P35 stays with D. For durable data I must also ensure C obtains the required current state; the placement calculation does not transfer bytes.”
Interviewer follow-up
What if B fails before a copy is made?
Reveal the follow-up answer
“Recovery needs another durable replica or retained history. A ring alone cannot reconstruct missing data.”
What the answer must demonstrate: Placement and durability are separate responsibilities.
“That changes many assignments at once, even though most existing machines are still healthy. Hash 35 changes remainder from 2 to 3, while 12 happens to stay at 0. Broad remapping can create expensive migration or cache misses; consistent hashing limits the affected ranges.”
Interviewer follow-up
Does modulo become impossible to use?
Reveal the follow-up answer
“No. It is simple for fixed membership or when managed logical buckets absorb physical changes. The issue is the resizing consequence.”
What the answer must demonstrate: Avoid claiming every modulo mapping necessarily changes.
“They give one physical host several separated ring positions, so it owns multiple smaller intervals. With a suitable distribution, this reduces random placement imbalance and can represent differing capacities. It adds token metadata and migration units; it does not create more independent machines.”
Interviewer follow-up
How do you place three replicas when several consecutive virtual tokens belong to one physical host?
Reveal the follow-up answer
“I walk eligible token positions but skip owners already selected, and enforce the required zone or rack diversity. Three tokens on one host are one failure domain, not three durable replicas. The token-placement rule and replica-placement policy are separate.”
What the answer must demonstrate: Count physical failure domains for replication.
One key P35 receives half of all reads. Will more virtual nodes split that hot key?
Reveal a model answer
“No. The same key still maps to one primary owner under this rule. I would consider read replication, caching, or request coalescing, while defining update and freshness behavior. Virtual positions improve distribution across many keys rather than splitting one indivisible key’s traffic.”
Interviewer follow-up
What other imbalance should you measure?
Reveal the follow-up answer
“Bytes per object. Equal key counts can hide one owner holding much larger values and exhausting storage first.”
What the answer must demonstrate: Key count, bytes, and traffic are different load measures.
How many of 1.2 million keys move when three balanced owners become four?
Reveal a model answer
“The expected share for the new equal-capacity owner is about one quarter, or 300,000 keys. I would label the balance and distribution assumptions. At 500 bytes each that is about 150 MB of payload before overhead, which helps estimate a controlled transfer.”
Interviewer follow-up
Why can a particular node insertion move a different fraction than that expectation?
Reveal the follow-up answer
“The estimate assumes balanced placements and a suitable key distribution. On a 0–99 ring, adding D40 between A20 and B50 moves (20,40], only 20% of that fixed space.”
What the answer must demonstrate: Qualify both arithmetic and assumptions.
A write updates P35 while its ownership moves from B to D. What must the migration protocol guarantee?
Reveal a model answer
“D needs a snapshot and the updates committed while that snapshot is copied. I would catch up, verify, and atomically change the authoritative routing generation under the migration protocol. B must forward or reject stale requests rather than keep an independent writable copy. After D accepts new writes, routing back to B requires reverse catch-up; retaining B’s old snapshot alone does not make rollback safe.”
Replication copies changes to additional replicas. Redundancy is the broader idea of spare resources: a spare machine, disk, or network link can be redundant without containing a usable data copy. Durability is the guarantee that a committed write survives a defined failure set. The replication protocol, durable storage, acknowledgment rule, and failover rules jointly determine that guarantee.
Suppose leader A acknowledges cart v41 before follower B receives it. Replication is configured, but permanently losing A can still lose that acknowledged write. Waiting for the required durable copies reduces this loss exposure while adding network/storage latency and making writes depend on those copies being reachable. Replication also copies a mistaken deletion, so it does not replace a backup.
“Received bytes may be only in memory. Durable bytes survive the specified storage failure model. Applied entries are visible to queries. B can have v41 durably logged while ordinary reads still show v40, so acknowledgment and read policy must account for different milestones.”
A leader acknowledges v41 before a follower receives it, then permanently fails. Explain the possible data loss.
Reveal a model answer
“A responds at .003, fails at .006, and B would receive the change at .008. If A’s storage is lost, the survivors have v40. I either accept that acknowledged-write loss window explicitly or wait for the required durable replica before answering.”
With three replicas and a one-replica-loss durability goal, why might the commit protocol wait for two durable copies instead of all three?
Reveal a model answer
“Two durable copies leave at least one copy of an acknowledged entry after any one participant is lost. With a safe election and commit protocol, the surviving majority preserves that committed history and can continue. Waiting for all three adds a copy but makes the slowest replica control acknowledgment and stops writes if any replica is unreachable. I would choose two only because it meets the stated one-failure contract; the count alone is not the safety proof.”
Interviewer follow-up
What if two simultaneous storage losses must be tolerated?
Reveal the follow-up answer
“I must revisit replica count, acknowledgment, and placement together. One surviving copy cannot preserve a write it never received.”
What the answer must demonstrate: Failure budget and acknowledgment must agree.
A write of v41 succeeds, but a subsequent session read returns v40. What should you inspect?
Reveal a model answer
“Check which replica answered and how far it had applied the write log. It may have saved v41 without making it readable yet. To read my own write, use the verified current leader or wait for a follower to apply the returned commit position. That position must still identify the right history after failover. A former leader or an arbitrary application version cannot prove freshness.”
Interviewer follow-up
Why is 100 ms of waiting insufficient?
Reveal the follow-up answer
“Lag is not bounded by that guess during overload or failure. I need evidence that the required update became visible.”
What the answer must demonstrate: Waiting a fixed time does not prove that the required update is visible.
“A shard owns a subset of records, while replicas store copies of that subset. Cart C17 can belong to one shard with three replicas. Adding shards can divide data and write work; adding followers preserves copies and can spread eligible reads. Each follower still has to process its shard’s write stream.”
Interviewer follow-up
Will ten followers give ten times the write capacity?
Reveal the follow-up answer
“Not by themselves. A single leader still orders the stream, and each follower must keep up with it.”
What the answer must demonstrate: Do not count duplicated processing as partitioned work.
A new leader B takes over from isolated leader A. What prevents A from continuing to commit writes?
Reveal a model answer
“A must lose the ability to commit new writes when B takes over. Missing heartbeats alone does not prove A stopped. Use the database’s safe election and fencing protocol to reject the old leader, then update routing so clients find B.”
Every replica contains a mistaken deletion. What next?
Reveal a model answer
“I stop the faulty job, restore retained history in isolation, identify C17’s last valid state, and verify the repair. Promoting another current replica cannot undo a deletion they all copied correctly. I would also check the full affected range.”
Interviewer follow-up
What proves that recovery plan works?
Reveal the follow-up answer
“A measured restore that validates records and application behavior within the objectives. Backup completion alone proves only part of the path.”
What the answer must demonstrate:Replicas and recovery history solve different failures.
What is the CAP theorem? Define C, A, and P, and explain the triangle with a concrete example.
Reveal a model answer
CAP says a distributed read/write system cannot guarantee both linearizableconsistency and completion of every request to a nonfailed participant when network partitions are allowed. C means clients observe one up-to-date copy: after a write completes, a later read must return it or a newer write. Formally, operations fit one valid history respecting real-time order. A means every such request eventually completes according to its contract. P means live replicas can be unable to exchange messages.
Draw C, A, and P at the triangle's vertices. Label CP as preserving one history while some operations wait or fail, AP as permitting completion with weaker consistency, and CA as requiring that partitions are excluded from the guarantee. Do not present P as a network failure you can disable in production.
For example, East and West both store S7 as free. They lose contact. East confirms client A's reservation. A later West read cannot learn that fact: returning free violates C; refusing or waiting without completion gives up A. The design should state which behavior is acceptable for that operation.
Interviewer follow-up
Why is “every replica has the same data at every instant” an inaccurate definition of CAP consistency?
Reveal the follow-up answer
Linearizability constrains observable operations, not instantaneous physical equality of every copy. A follower can lag if the system routes, waits for, or validates reads so completed operations still fit one legal real-time order. A read overlapping a write may legally appear before or after it. But if the write completed before the read began, an older value is invalid in the absence of an intervening write. The physical replication and the visible consistency promise are different levels.
What the answer must demonstrate: State the theorem before the caveats; define all three letters and use one completed-write/later-read partition trace.
“The client reached a working participant but did not complete the requested seat read. The server replied quickly, which is useful operationally, but refused the object operation. I would count that separately from a valid ‘already reserved’ result and separately from the product’s latency target.”
Interviewer follow-up
Does returning ‘already reserved’ sacrifice availability?
Reveal the follow-up answer
“Not when that is a valid result established by the reservation operation. It reports a business outcome. Inventing that result without authority just to avoid an error would violate the operation’s contract.”
What the answer must demonstrate: Separate infrastructure failure from legitimate business rejection.
Can a partition happen while both databases are healthy?
Reveal a model answer
“Yes. East and West may both run normally and answer their local clients while network messages between them are dropped. That is why checking each process’s health is insufficient. I need to know which communication and authority assumptions an operation requires.”
“It reduces the chance of losing communication, but cannot prove communication will always work. I still define behavior for the residual case where every usable path fails.”
What the answer must demonstrate: A network partition is not necessarily a server crash.
East and West start with S7 free, then become partitioned. East confirms a reservation at 10:00:02; a West read begins at 10:00:03. Why can West not guarantee a linearizable answer while completing every such read?
Reveal a model answer
“West has the same local state in several possible histories: client A reserved in East, someone else reserved, or nobody wrote. No East message has arrived. Its old null value cannot distinguish them. Answering immediately may choose the wrong history; waiting for information can prevent completion during a continuing partition.”
Interviewer follow-up
Could synchronized clocks reveal the missing write?
Reveal the follow-up answer
“Clocks can tell West that time passed, but not who wrote or whether a write happened. Timing assumptions may support particular protocols, but time alone does not carry the missing data.”
What the answer must demonstrate: Explain the missing information, not just repeat ‘choose two.’
How would you handle the last seat during a partition?
Reveal a model answer
“I would allow only the participant with valid write authority to perform the atomic available-to-reserved transition. A disconnected minority would decline it. That may stop some purchases, but a successful confirmation then means the seat was reserved by the node currently authorized to make that decision. I would specify the quorum and safe leader change rather than relying on a product label.”
“If safe progress requires both, a split leaves neither side able to complete new writes. I would discuss a third voting participant and failure-domain placement, or accept the two-node availability cost.”
What the answer must demonstrate: Adding replicas is not the same as defining a safe election protocol.
Does preventing double sales imply every read is CAP-consistent?
Reveal a model answer
“No. I can send all reservations through one atomic authority while serving a stale seating map elsewhere. The business invariant can hold even when that display is not linearizable. Conversely, a correctly ordered store can still oversell if my application uses an unsafe read-then-write algorithm.”
Interviewer follow-up
How would you fix that unsafe algorithm?
Reveal the follow-up answer
“Make checking availability and assigning the owner one protected operation, using a conditional update or suitable transaction. Read freshness alone does not make two separate operations atomic.”
What the answer must demonstrate:CAP C and application invariants are related design concerns, not identical definitions.
A reservation request times out without a known outcome. How should the client retry?
Reveal a model answer
“Reuse the operation identifier and ask the authority for the durable outcome. A timeout means the response was not received; it does not prove the reservation failed. If the old attempt committed, return that result. If it did not, process the retry under the same ownership rules.”
Interviewer follow-up
Should West create a new reservation while East is unreachable?
Reveal the follow-up answer
“Only if West can safely take responsibility for the reservation. If East may already have reserved the seat, creating an unrelated reservation at West could create two conflicting bookings.”
What the answer must demonstrate: A missing response is an unknown outcome.
“Replicas must converge on the protocol’s authoritative history, and obsolete writers must remain fenced. I would verify catch-up before routing reads that promise current state. If our policy allowed conflicting writes, I also need an explicit business repair policy; network recovery alone cannot choose who deserves a promised seat.”
Interviewer follow-up
Can the whole site have one useful AP or CP label?
Reveal the follow-up answer
“Only as shorthand for a specified operation and failure model. A stale advisory map and an authoritative reservation already make different choices. AP also does not define eventual convergence or conflict resolution; I must explain how accepted updates propagate and reconcile after communication returns.”
What the answer must demonstrate: Recovery must honor promises made before and during the fault.
What is a consistency model? Explain it using a write of version 11 followed by a read.
Reveal a model answer
A consistency model defines the read results and operation orders a system allows. If client A completes a write of version 11 and client B then reads, linearizability forbids the old version 10 when no other write intervened. Eventual consistency may temporarily allow version 10. The choice describes a visible contract, not whether the title text is factually correct.
I ask which operation and scope need the guarantee. A title read after a completed save suggests linearizability for that object. Updating title and ownership together also requires a transaction contract. I describe one forbidden history before selecting a database.
What the answer must demonstrate: Define permitted observations and the object or transaction scope.
An interviewer says “the system must be consistent.” Which meaning should you clarify?
Reveal a model answer
I ask whether the requirement concerns read visibility or a business invariant. For CAP consistency, a write that completes before a read starts must be visible to that read, or superseded by a newer write. More generally, I name the required consistency model, such as linearizable or causal. ACID consistency means transactions preserve rules such as nonnegative stock. I would state the operation and show a concrete forbidden result.
No. The service may route a read to an authoritative copy or wait until it can satisfy the guarantee. A stale successful read after a completed write violates linearizability when no later write explains it; a lagging physical copy alone does not. Refusing or indefinitely waiting for an affected operation sacrifices CAP availability.
What the answer must demonstrate: Connect the familiar current-value explanation to the formal model, and keep ACID validity separate.
A write from v10 to v11 overlaps a read on another client. Must a linearizable read return v11?
Reveal a model answer
“Not necessarily. Under linearizability the read may take effect before or after the concurrent write. I would inspect invocation and response intervals; a read beginning after the write completed is the clearer test.”
Interviewer follow-up
Can the server always return the old value while calling every write concurrent?
Reveal the follow-up answer
“No. The recorded operation intervals constrain that explanation, and completed earlier writes must be respected.”
What the answer must demonstrate: Do not replace the definition with a vague latest-value rule.
“Client A completes writing v11, then an independent client B starts a read and gets v10. With no other operations, a total order can put client B’s read first, preserving each client’s order. Real-time completion forbids that placement under linearizability.”
Interviewer follow-up
What if the writer performs that later read in the same session?
Reveal the follow-up answer
“With no intervening writer, returning v10 would violate its own write-then-read order, so that history is not sequentially consistent either.”
What the answer must demonstrate: Keep process order separate from wall-clock order.
How do you stop replies appearing before their comments?
Reveal a model answer
“I attach the parent or a sufficient dependency context to client B’s reply. A replica cannot expose the reply until it can expose that history. That is a visibility rule, not just sorting by arrival timestamp.”
Interviewer follow-up
Must two unrelated comments have the same order everywhere?
Reveal the follow-up answer
“Causal consistency does not require that. If the product needs one conversation sequence, I add an ordering mechanism and accept its cost.”
What the answer must demonstrate: Dependencies do not imply a total order for independent writes.
How can a session preserve read-your-writes when failing over from a replica at v11 to one at v10?
Reveal a model answer
“The save response carries a storage position or version context. The next server must prove it has applied that context before answering. If it cannot, it routes or waits; silently returning v10 violates the session promise.”
Interviewer follow-up
Is sticky routing enough?
Reveal the follow-up answer
“It helps during normal operation, but cannot preserve the promise when the pinned server fails and the replacement is behind.”
What the answer must demonstrate: Describe failover as well as the normal request path.
Does a five-second TTL guarantee data no older than five seconds?
Reveal a model answer
“Only under additional assumptions. If a cache fills from a replica already thirty seconds behind, a fresh cache entry is still stale. I need an authoritative reference, propagation limits, and behavior when the bound cannot be met.”
Interviewer follow-up
What would you measure?
Reveal the follow-up answer
“I would measure source-version age or replication lag along the entire read path, with clock assumptions made explicit for time-based bounds.”
What the answer must demonstrate:Cache age and source age differ.
“No. It tells us which edits depend on which earlier edits. Independent edits still need a conflict policy, such as preserving both versions for the user or a domain-specific merge. A last-writer rule chooses a winner but can lose intent.”
Interviewer follow-up
Would a timestamp winner always identify the last human edit?
Reveal the follow-up answer
“No. Clock error and concurrent work make that claim unsafe; a timestamp can define an arbitration rule without representing human intent.”
What the answer must demonstrate: Separate causal ordering, convergence, and application semantics.
Why not require linearizability for every read in a collaborative application?
Reveal a model answer
“It may be acceptable, especially at modest scale, but I would compare the added coordination latency and the operations that may become unavailable with what the product requires. The editing session and comment dependencies can often have clear weaker contracts, while ownership still needs stricter enforcement.”
Interviewer follow-up
What must you avoid when mixing guarantees?
Reveal the follow-up answer
“I must prevent a weaker cache or replica path from serving an operation whose security or correctness contract is stronger.”
What the answer must demonstrate: Make the choice per operation, not per marketing category.
Atomicity makes a transaction’s changes commit together or abort together. Isolation controls how concurrent transactions observe and interfere with one another. In the roster example, two atomic leave requests can both read 2 and update different rows, leaving 0 on duty under snapshot isolation. The database needs a concurrency rule that protects the shared business condition.
Interviewer follow-up
Would storing both rows on one database server remove the concurrency race?
Reveal the follow-up answer
No. One database server can execute concurrent transactions. I still need an appropriate isolation level, a constraint, or a shared guard acquired before the decision’s read.
What the answer must demonstrate: Separate all-or-nothing changes from safe concurrent decisions.
Explain nonrepeatable and phantom reads without jargon.
Reveal a model answer
“A nonrepeatable read occurs when transaction T1 reads row D2 twice and observes different committed values because another transaction changed it. A phantom occurs when T1 repeats a predicate query, such as all on-duty rows, and sees a new or missing matching row. The first concerns an existing row’s value; the second concerns membership in a result set.”
Interviewer follow-up
Does locking the rows currently matching a query prevent a new matching row from being inserted?
Reveal the follow-up answer
“Not generally. Protecting the queried set may require predicate or range protection, or an agreed guard row.”
What the answer must demonstrate: Use one row versus a matching set.
“I would ask which database. The SQL standard’s minimum guarantees allow that phenomenon, while PostgreSQL Repeatable Read uses a stable snapshot and prevents it. Neither statement means PostgreSQL Repeatable Read prevents our write-skew example.”
Interviewer follow-up
Why can a stable snapshot still be dangerous?
Reveal the follow-up answer
“Two transactions can make incompatible decisions from it and write different rows without a same-row conflict.”
What the answer must demonstrate: Do not generalize product behavior from the level name.
Two transactions read revision 8 and both assign 9. How do you prevent this lost increment?
Reveal a model answer
“I use an atomic increment or a compare-and-update against the expected revision, checking whether it succeeded. Reading 8 in application code and later assigning 9 in both requests loses one increment.”
Interviewer follow-up
Does fixing a revision counter automatically protect a separate multi-row count constraint?
Reveal the follow-up answer
“Only if the revised protocol actually uses the shared row to serialize or validate the entire decision. A separate counter fix alone does not.”
What the answer must demonstrate: A local race fix must cover the business decision to enforce it.
To protect count(on_duty) >= 1 using a shared guard row, when must the guard be locked relative to reading the count?
Reveal a model answer
“Before reading the state used to decide whether someone may leave. I use a transaction pattern whose post-lock query observes the previous holder’s committed result; with Read Committed, a subsequent query gets a fresh statement snapshot.”
Interviewer follow-up
What if the transaction already read its snapshot before waiting?
Reveal the follow-up answer
“I cannot assume acquiring a lock refreshes that earlier snapshot. I must restart or use an isolation-specific safe pattern.”
What the answer must demonstrate: Lock timing and snapshot timing must agree.
T2 receives a serialization failure. What does the application do?
Reveal a model answer
“Abort the failed attempt and retry the complete transaction: reads, validation, and writes. If another transaction reduced the on-duty count to one, the new execution must reject the off-duty transition. Retrying only the final write reuses an invalid decision.”
Interviewer follow-up
Why not resend only the UPDATE?
Reveal the follow-up answer
“That repeats the write while discarding the validation the transaction was supposed to protect.”
What the answer must demonstrate: Retries must recompute the decision.
How do you emit an off-duty notification only for a committed transition when its transaction may abort and retry?
Reveal a model answer
“I record the notification intent atomically with the successful roster transaction. A separate worker sends it using a stable event identifier. The retried transaction body must not perform irreversible external work.”
Interviewer follow-up
What happens if the worker sends the message and crashes before acknowledging?
Reveal the follow-up answer
“Delivery may repeat, so the receiver or publication mechanism needs deduplication where required. The outbox closes the database-to-event gap, not every downstream effect.”
What the answer must demonstrate: Explain which database changes commit together and which later message delivery still needs deduplication.
What operational costs should you measure for a guard row that serializes all changes to one roster?
Reveal a model answer
“I measure wait time, transaction length and contention by roster. A long-held guard is a latency bottleneck even if CPU looks idle. I keep the protected work short and test simultaneous leave, deletion and transfer operations.”
Interviewer follow-up
When would you change the design?
Reveal the follow-up answer
“If one roster becomes a hot coordination point or workflows span many rosters, I would revisit the invariant’s ownership and transaction scope instead of simply increasing connection count.”
What the answer must demonstrate: More concurrency can worsen a serialized bottleneck.
A quorum is a required response set, such as two of three controllers. Consensus makes those controllers agree on an ownership decision or committed log despite the failures it tolerates. A lease gives ownership for a limited interval. Fencing adds an increasing ownership token that the output store checks atomically with a write.
For export E9, controllers grant W1 epoch 7. W1 pauses; the lease expires; controllers agree to grant W2 epoch 8. W2 publishes with token 8. If W1 resumes and presents 7, the store rejects it. The lease did not stop W1's CPU from executing; the fencing check stops its stale effect after newer authority reaches the resource. Quorum overlap helps the agreement proof but does not supply a complete consensus protocol.
Interviewer follow-up
What changes if R=1 and W=1?
Reveal the follow-up answer
“The groups may be R1 and R3, with no common member. A read can entirely miss an acknowledged write.”
What the answer must demonstrate: Start from actual sets rather than a memorized equation.
“With three replicas, read groups {R1,R2} and {R2,R3} do overlap at R2. But suppose only R1 saw an incomplete write of v9 while R2 and R3 still have v8. A first read returns v9 from R1, then a later read through R2 and R3 returns v8 without another write. The problem is that the read exposed a value without preserving it for later reads. Quorum intersection alone does not define safe version selection, write-back, commitment, or recovery.”
Interviewer follow-up
Can a clock timestamp settle it?
Reveal the follow-up answer
“Not by itself. Clocks may disagree, and a high timestamp does not prove an update belongs to the committed history.”
What the answer must demonstrate: Use overlapping replica sets and non-overlapping-in-time reads; explain why a selected value must remain visible to later reads.
What does consensus provide when three controllers assign one owner for export job E9?
Reveal a model answer
“It gives the controllers one agreed sequence of ownership transitions, so W1 expiry and W2’s epoch-8 grant are not independently invented on different copies. I would use a proven replicated-log protocol whose election and commit rules preserve the history after controller failure.”
Interviewer follow-up
Does that protocol automatically publish the export once?
Reveal the follow-up answer
“No. Agreement commits the ownership metadata. The output service must enforce ownership and deduplicate publication at its own boundary; otherwise two workers may still produce conflicting external effects.”
What the answer must demonstrate: Agreement on metadata does not atomically include every external effect.
What happens when two of three controllers are unreachable?
Reveal a model answer
“Only one remains, so the majority protocol cannot safely advance ownership. I would stop new grants and report reduced availability. I would not let the isolated replica infer that its stale state is now authoritative because it is the only one this client can reach.”
Interviewer follow-up
Would four or five controllers improve the number of unavailable controllers tolerated?
Reveal the follow-up answer
“Four voters need three votes, so they still tolerate only one unavailable voter. Five need three and can tolerate two, provided the remaining three communicate and satisfy the protocol’s progress requirements. The benefit comes from the voting threshold and failure placement, not simply a larger count.”
What the answer must demonstrate: Distinguish safety from continued progress.
W1 has fencing token 7; replacement W2 publishes with token 8. W1 resumes. What must the output store check?
Reveal a model answer
“W2’s publication has fence 8, so the output store has atomically recorded that generation with the manifest. W1 arrives carrying 7. The store rejects 7 before changing the protected state, preventing W1 from replacing W2’s newer result.”
Interviewer follow-up
What if W1 checks the lock before making a separate write?
Reveal the follow-up answer
“Ownership can change after the lock check but before the write. Check the fencing token, update the manifest and save the accepted token together in one atomic operation. Saving the token also prevents a restart from forgetting which worker is current.”
What the answer must demonstrate: A separate preflight check leaves a race.
Does a fence instantly revoke old work everywhere?
Reveal a model answer
“Not necessarily. A resource comparing against its latest accepted fence learns about generation 8 when that newer authority reaches it. It prevents older writes after that point. If the requirement is immediate revocation everywhere, I need current-ownership validation or another stronger coordinated boundary.”
Interviewer follow-up
Could W1’s computation continue harmlessly?
Reveal the follow-up answer
“Yes, if its consequential publication is prevented. Wasted computation and an unauthorized state change are different concerns.”
What the answer must demonstrate: Describe the precise fencing guarantee rather than implying physical process termination.
A valid worker retries publication with the same fencing epoch 8. Why is an operation idempotency key still needed?
Reveal a model answer
“Epoch 8 says W2 is an eligible owner. It does not distinguish one publication attempt from a retransmission of the same attempt. I use a stable publication ID so a lost response does not create duplicate effects while that ownership is still valid.”
Interviewer follow-up
Can an old epoch with a new operation ID be accepted?
Reveal the follow-up answer
“No. Idempotency is not permission. It must still pass the ownership check.”
What the answer must demonstrate: Operation identity and authorization are independent checks.
What do you reconcile after the controller outage ends?
Reveal a model answer
“I inspect the committed ownership history, the output store’s accepted fence and manifest, and any unfinished files. A worker saying it finished is weaker than the protected publication record. I then resume or retry with stable IDs and valid ownership rather than blindly rerunning every reported job.”
Interviewer follow-up
Can a lagging controller issue grants while it catches up?
Reveal the follow-up answer
“Not under our chosen current-authority contract. It must participate according to the consensus protocol before serving authoritative ownership decisions.”
What the answer must demonstrate: Recovery must consult the state that actually governs the external result.
Idempotency means repeating the same logical operation has the same intended effect as doing it once. A timeout does not prove failure: O17 may have committed before its response was lost. U9 retries buy-204 with the same caller and payload so the service returns O17 instead of making another purchase.
Interviewer follow-up
Does idempotency require the response bytes and every network message to be identical?
Reveal the follow-up answer
No. It concerns the intended effect. Retries can produce additional network messages or different status details while still referring to the same one business operation.
What the answer must demonstrate: Distinguish one business effect from one transport attempt.
“One caller’s logical intent, such as U9’s purchase attempt buy-204. I scope it to the authenticated caller and compare a canonical payload fingerprint. A fresh network retry reuses it; a new intended purchase gets a different key.”
Interviewer follow-up
What if the payload changes under the same key?
Reveal the follow-up answer
Reject a changed operation, target or canonical payload under the same scoped key. The stored result still needs current authorization; knowledge of the key is not permission to inspect another resource.
What the answer must demonstrate: Separate caller, intent, and payload.
Two requests both see no saved result. How is one order guaranteed?
Reveal a model answer
“The claim and business effect must share an atomic boundary, such as a unique request record and order insert in one transaction. A separate check-then-insert can let both proceed. The losing concurrent attempt waits for or retrieves the winner’s outcome.”
Interviewer follow-up
What if a long-running task is in progress?
Reveal the follow-up answer
Return its status or wait up to a limit. If a new worker takes over, give it a new ownership version and atomically reject old versions when saving local results. For external calls already sent, use the receiving service’s duplicate protection or check their outcomes.
What the answer must demonstrate: Show the atomic boundary.
“Only if the contract prevents valid retries after that minute or another durable identity prevents repetition. Deleting the record can make a delayed duplicate look like a new operation. I align retention with the retry horizon, business identifiers, and downstream retention.”
Interviewer follow-up
What if the provider retains keys for less time than we do?
Reveal the follow-up answer
Our workflow must stop blind retries outside the provider guarantee and reconcile through a durable provider resource ID or another supported status path.
What the answer must demonstrate: Treat deduplication retention as part of correctness.
Why doesn’t a local transaction make the external charge exactly once?
Reveal a model answer
“The provider is outside that transaction. It can charge and lose its reply before we save the result. I use one durable provider attempt key, query or reconcile its outcome, and process duplicate notifications safely. The local order and remote charge have separate commit boundaries.”
Interviewer follow-up
What if the provider lacks those capabilities?
Reveal the follow-up answer
I cannot invent the guarantee. I would state the residual uncertainty and design reconciliation, compensation, or an operational resolution path.
What the answer must demonstrate: Avoid blanket exactly-once claims.
Three layers each make three attempts. What reaches the bottom?
Reveal a model answer
“In the worst simple nesting, up to 27 calls for one user operation. That extra work can keep a struggling dependency down. I choose one retry layer or a shared budget, cap attempts and total time, and stop retrying permanent failures.”
How do timeouts relate to a two-second user budget?
Reveal a model answer
“The two-second budget becomes an absolute deadline two seconds after the request starts. Each downstream call gets at most the remaining time as its timeout, including planned retries. Independent two-second waits at every layer can greatly exceed it and keep doing work after the user has left.”
Interviewer follow-up
Does cancellation undo a completed action?
Reveal the follow-up answer
No. Cancellation can stop unnecessary pending work, but committed effects still need normal reconciliation or compensation.
What the answer must demonstrate: Distinguish stopping work from reversing it.
How do you stop a slow export dependency from taking down checkout?
Reveal a model answer
“I separate concurrency pools so exports cannot consume all checkout workers or connections. I bound queues and use deadlines, then reject or defer lower-priority work when capacity is exhausted. A breaker can limit calls to the unhealthy dependency while probing recovery.”
It accepts work without a credible completion time and can exhaust resources. Availability must include a meaningful service contract, not merely enqueueing forever.
What the answer must demonstrate: Protect a finite resource and explain overload behavior.
A producer submits a message; a broker stores and delivers it; a consumer processes it. A message queue buffers that work so its execution can happen asynchronously, after the initial request has durably accepted it. Accepted and completed are different states. A photo upload can return a job ID while its thumbnail is still pending.
Backpressure limits incoming or concurrent work when downstream processing cannot keep up. If arrival is 600 jobs/s and workers can process 400 jobs/s, the backlog grows by 200 jobs/s. Bound admission or increase processing capacity instead of treating an unbounded queue as a solution. For reliable handoff, an outbox can commit the photo and job intention together; consumers still need duplicate-safe effects because delivery or acknowledgment can repeat.
Interviewer follow-up
Why not keep the HTTP request open until completion?
Reveal the follow-up answer
“That may be reasonable for short bounded work, but variable processing and bursts tie up request capacity and make retries harder. Our asynchronous contract separates those concerns.”
What the answer must demonstrate: Acceptance and completion are different promises.
How does a work queue differ from a retained event log?
Reveal a model answer
“The work queue assigns J501 and manages retry eligibility. A retained log lets consumers replay an ordered sequence from their positions, possibly in independent groups. I would choose based on work assignment versus replay needs and inspect the real retention and delivery semantics.”
A database commits photo P501, then the process dies before publishing its job J501. How does a transactional outbox close that gap?
Reveal a model answer
“I store the photo and an outbox intention in one transaction. A relay can find committed unpublished intentions after recovery. This prevents P501 being accepted without a durable record that processing must happen, while avoiding a false claim of one transaction across database and broker.”
Interviewer follow-up
Could the relay still publish twice?
Reveal the follow-up answer
“Yes. It might publish and crash before marking completion, so workers must recognize duplicate J501 deliveries.”
What the answer must demonstrate:Outbox solves a missing handoff, not every duplicate.
The worker commits the result and crashes before acknowledging. Walk the retry.
Reveal a model answer
“The job is delivered again after the broker did not record completion. A worker reads the durable job receipt and returns the established outcome. If two duplicate workers race before either receipt exists, a unique job-ID constraint and one result transaction choose the winner; the loser rolls back and reads that winner’s outcome. A preflight receipt lookup alone would not prevent two concurrent effects.”
Interviewer follow-up
What if the receipt was written before the effect?
Reveal the follow-up answer
“A crash could make later workers skip work that never completed. Evidence of completion must be committed with the recoverable effect.”
What the answer must demonstrate:Deduplication placement determines correctness.
Does a consumer’s local deduplication receipt make an external API call exactly once?
Reveal a model answer
“No. A remote object write is outside the result database transaction. I give each render attempt an immutable object key and use the stable photo/version/recipe identity for the logical job. The database transaction chooses one reference and records its receipt; it never lets a duplicate overwrite the chosen bytes. Cleanup cannot delete an active attempt that may still publish. A different provider effect needs that provider’s idempotency contract or reconciliation of unknown outcomes.”
Interviewer follow-up
Can you just write ‘done’ before the remote call?
Reveal the follow-up answer
“No. If the process dies after recording done but before the remote call, recovery can suppress the only attempt. I need a durable state machine that distinguishes intent, an uncertain external outcome, and confirmed completion, plus a safe retry or reconciliation path.”
What the answer must demonstrate: Local atomicity does not automatically include a remote effect.
A thumbnail job for photo version 2 finishes after version 3 is published. What should the result commit check?
Reveal a model answer
“The authoritative ready update checks which photo version it belongs to. A stale version-2 result may be retained or cleaned up, but it cannot replace the version-3 reference. Ordering events by photo can help, yet the conditional update protects against retry and completion reordering.”
Interviewer follow-up
Can J501 be reused for different image bytes?
Reveal the follow-up answer
“Not silently. I would bind it to the request identity/version and reject conflicting reuse or create a distinct job.”
What the answer must demonstrate: Stable identity must represent stable intent.
For 60 seconds, arrivals are 600 jobs/s and workers complete 400/s. Arrivals then fall to 200/s. Calculate backlog growth and drain time.
Reveal a model answer
“Six hundred arrivals minus four hundred completions gives two hundred extra jobs per second. Over sixty seconds that is twelve thousand jobs. When arrivals fall to two hundred, the spare capacity is two hundred, so recovery takes about sixty seconds under stable-rate assumptions.”
Interviewer follow-up
What if arrivals remain at four hundred?
Reveal the follow-up answer
“There is no spare capacity, so the accumulated backlog remains. I need extra processing capacity or lower arrivals to reduce it.”
What the answer must demonstrate: Use service rate minus arrival rate when estimating drain time.
J501 contains an image that can never be decoded. Should it retry forever?
Reveal a model answer
“No. I would classify the permanent failure, stop after a bounded policy, store diagnostics, and expose a failed state or dead-letter workflow. Infinite retries consume capacity and can delay valid work. Any operator replay should be controlled and preserve job identity and version rules.”
Interviewer follow-up
Which alert reflects the client’s experience most directly?
Reveal the follow-up answer
“Oldest pending-job age or completion-latency violations, alongside failure rate. Queue length alone does not tell me how long the client’s job has waited.”
What the answer must demonstrate: Bound both retry effort and user-visible delay.
A distributed transaction spans multiple transactional participants and needs a coordinated commit-or-abort outcome. A saga addresses a related business need through separately committed local transactions and compensation. For O81, creating an order, reserving 2 mugs, and authorizing $24 can succeed or fail separately. A capable 2PC system coordinates one commit decision; a saga records local progress and compensates failures. I first ask whether the work could remain in one simpler database transaction.
When the invariant and data already fit one database ownership boundary. A saga adds visible intermediate states and recovery work. Putting separate databases in one application process does not make them one transaction domain.
What the answer must demonstrate: Identify the actual independent commit boundaries.
“The participant has prepared enough durable state and retained the necessary protections to honor a later commit decision. It is stronger than saying the request looks valid right now.”
Interviewer follow-up
Why can coordinator failure block progress?
Reveal the follow-up answer
“A yes voter may not know whether commit was already decided, so it cannot safely invent an abort solely from a timeout.”
What the answer must demonstrate: Prepared is a durable protocol state, not a best-effort check.
“2PC coordinates the final commit or abort outcome. Isolation depends on the concurrency-control protocol over the affected reads and writes. I would not claim serializability just because every participant votes on one decision.”
Interviewer follow-up
What would you inspect?
Reveal the follow-up answer
“I would inspect locking or validation across participants and show whether concurrent transactions admit a valid serial order.”
What the answer must demonstrate: Atomic commit and isolation solve different parts of correctness.
A payment authorization A81 times out with no known result. What should a durable workflow do next?
Reveal a model answer
“Save the outcome as unknown and use A81 to check with the provider. I do not create A82 just to retry: A81 may already have succeeded.”
Interviewer follow-up
What if the provider has neither idempotency nor a status query?
Reveal the follow-up answer
“The ambiguity cannot be eliminated by our local database alone. I need another provider-supported reconciliation mechanism or a product process that explicitly handles unresolved outcomes.”
What the answer must demonstrate: Do not promise exactly-once effects across an unsupported boundary.
An inventory hold expires at 120 seconds and payment authorization succeeds at 125. Can the order be confirmed?
Reveal a model answer
“Not from the authorization alone. Inventory must atomically verify or convert a valid hold, and H81 is expired. I keep confirmation conditional and void the authorization while cancelling the order.”
Interviewer follow-up
What if a callback races with cancellation?
Reveal the follow-up answer
“Both transitions check the durable workflow state, and inventory independently checks the allocation condition. A late callback cannot overwrite a terminal cancellation.”
What the answer must demonstrate: Two authorities must enforce their own conditions.
“The cancellation has an outstanding cleanup state with a stable void identifier. A worker retries or queries it, and an age-based alert exposes work that cannot finish automatically.”
No. The order can be blocked from fulfillment while authorization cleanup remains pending. A timeout does not prove the void failed, and retrying after the provider’s deduplication window expires may create a different external effect.
What the answer must demonstrate: Do not hide unfinished compensation behind a terminal label.
“It atomically records the local state transition and the intent to send the next message. After a crash, the relay can find that intent. The relay may publish twice, so consumers still need idempotent handling.”
Interviewer follow-up
Does it atomically commit the provider’s authorization?
Reveal the follow-up answer
“No. That remains a remote effect whose uncertain outcome must be reconciled separately.”
What the answer must demonstrate: Keep the outbox guarantee within its actual transaction boundary.
Would you use orchestration or choreography for an order workflow with inventory holds, payment authorization, and compensation?
Reveal a model answer
“I would start with an explicit coordinator because the order’s deadlines, compensation and user-visible status form one workflow that operators must inspect. Services still own inventory and authorization details.”
Interviewer follow-up
When could choreography fit?
Reveal the follow-up answer
“A few independent reactions to a completed fact, such as analytics and notification, may work well as subscribers. I would still assign ownership for failures and avoid an implicit cycle of events nobody can reconstruct.”
What the answer must demonstrate: Explain operational ownership instead of declaring one style universally better.
For chat, define a freshness target such as new messages normally appearing within 300 ms; this is an illustrative product target, not a property automatically guaranteed by a transport. Short polling repeats a request on a timer, creating idle traffic and up to roughly one interval of waiting. Long polling holds a request until data arrives or it times out, then the client starts another. SSE keeps an HTTP response open for text events from server to client. WebSocket maintains a full-duplex framed message channel.
For infrequent notifications, polling may be sufficient. For mostly one-way live updates, SSE plus ordinary HTTP commands can be simple. For frequent chat messages, typing, and acknowledgments in both directions, WebSocket is a reasonable choice. All choices need authentication, bounded buffering, reconnect, and a durable cursor/history policy; the socket alone cannot restore missed messages.
100,000 clients poll every five seconds. An event arrives at :02 between polls at :00 and :05. Estimate idle QPS and event delay.
Reveal a model answer
“100,000 clients divided by a five-second interval produce 20,000 requests/s even without updates. An event at :02 waits three seconds until the :05 poll, plus network and processing time. Uniformly timed arrivals wait roughly half an interval on average.”
Interviewer follow-up
How would you reduce load?
Reveal the follow-up answer
“Increase the interval, reduce active polling when appropriate, or change the transport. A longer interval has a clear freshness cost.”
What the answer must demonstrate: State the arrival and interval assumptions.
“The server holds one HTTP request until an update exists or the timeout expires. For example, a request after cursor 500 returns event 501, and the client immediately requests after its applied cursor again. The response may contain a batch; long polling means one response per request, not necessarily one event. History bridges the short gap before the next held request.”
Interviewer follow-up
Can a message arrive during the reconnect gap?
Reveal the follow-up answer
“Yes. The next request asks after the known ID, so durable history bridges the gap rather than relying on perfect timing.”
What the answer must demonstrate: Explain wait, response, reissue, and timeout.
“For HTTP/1.1, the client requests an upgrade and the server validates it and returns 101 before exchanging WebSocket frames. HTTP/2 and HTTP/3 have extended-CONNECT mechanisms when supported. I would specify what our gateway and clients actually support, authenticate the session, validate browser Origin, and authorize subscriptions; protocol negotiation alone grants no user permission.”
Interviewer follow-up
Does a successful handshake guarantee message persistence?
Reveal the follow-up answer
“No. It establishes the channel. The application still needs a durable history and an acknowledgment/resume contract.”
What the answer must demonstrate: Handshake, authentication, and durability are distinct mechanisms.
Could the receiving client send messages while receiving SSE?
Reveal a model answer
“Yes. The receiving client can receive a continuing event-stream response and send commands through separate HTTP POST requests. SSE is one-way on that stream, not a prohibition on the browser making other requests. It is attractive when live traffic is primarily server-to-client.”
Interviewer follow-up
What happens to binary attachments?
Reveal the follow-up answer
“I would normally upload and retrieve them through a separate media path, using events to carry metadata or references.”
What the answer must demonstrate: One-way stream does not mean one-way application.
A client receives event 501 but reconnects with last-applied cursor 500. How should replay work?
Reveal a model answer
“Replay event 501 from durable history and apply it idempotently by message ID. Cursor 500 must mean the application applied every event through that position in the relevant stream. Native SSELast-Event-ID can advance before the handler durably applies an event, so I would use the explicit application cursor for this stronger replay contract. A gateway writing bytes is not evidence that the recipient recorded the update.”
Interviewer follow-up
How do you avoid losing events between replaying history and switching to live delivery?
Reveal the follow-up answer
“Register a bounded live buffer first, record the latest committed event position H, replay after the client cursor through H, then deliver buffered events after H and continue live. Deduplicate overlap and preserve stream order. If history expired or the buffer overflows, require an explicit resynchronization instead of silently skipping the gap.”
What the answer must demonstrate: Connection delivery and application progress can differ.
How should a gateway handle a receiving client that consumes events slower than they arrive?
Reveal a model answer
“I bound the outgoing buffer. Depending on the event contract, I can drop optional typing updates, summarize state, or disconnect and resume durable messages later. I cannot let one slow client grow gateway memory without limit.”
Interviewer follow-up
Can I drop an undelivered durable message silently?
Reveal the follow-up answer
“Not if the product promised recoverable delivery. It must remain available through history and the resume path, or the client must be told the gap cannot be recovered.”
What the answer must demonstrate: Separate replaceable hints from durable events.
Why can a simultaneous reconnect after a gateway outage cause another outage?
Reveal a model answer
“A large connection outage can cause every client to reconnect and replay simultaneously. I would use randomized retry delays, admission control, and bounded replay work while protecting the history store. A healthy gateway fleet can still overload its shared dependencies during recovery.”
Interviewer follow-up
What access check happens on reconnect?
Reveal the follow-up answer
“Reauthenticate and reauthorize the requested subscriptions, including any changes while the client was offline.”
What the answer must demonstrate: Recovery traffic and permission changes are part of the protocol.
It uses randomization to obtain a useful space or performance tradeoff. Some probabilistic structures answer exactly; the Bloom filter is an approximate membership summary with a defined error model. In the Bloom example, A sets bits 2 and 7 and B sets 7 and 12. C tests 2 and 12, so the filter says possibly present even though C was never inserted: a false positive. It no longer knows which item set each bit.
Interviewer follow-up
What should a crawler do with that positive result?
Reveal the follow-up answer
Check the exact URL set. Skipping C solely because of a Bloom positive can permanently omit a new page. A negative saves a preliminary lookup only under the filter’s coverage assumptions; the exact unique insert still handles concurrent claims.
What the answer must demonstrate: Name the supported question, error direction, and business consequence.
Under what coverage and update assumptions is a Bloom-filter negative safe to trust?
Reveal a model answer
“It proves absence from a correctly maintained filter’s inserted set. To infer absence from the database, the filter must cover that database state. A stale or interrupted rebuild may omit real entries.”
Interviewer follow-up
How do you survive an incomplete filter?
Reveal the follow-up answer
“Bypass the filter, or trust it only for data its coverage record proves complete. Keep the exact atomic database claim when scheduling a URL.”
What the answer must demonstrate: State which set the guarantee describes.
Estimate memory for one million URLs at 1% false positives.
Reveal a model answer
“Using the standard idealized formulas, I need about 9.59 million bits, or 1.20 decimal MB, and about seven hash positions per item. I would add implementation overhead and headroom for growth.”
Interviewer follow-up
What happens at two million entries without resizing?
Reveal the follow-up answer
“More bits are set and false positives rise; the original 1% target no longer holds.”
What the answer must demonstrate: Keep bits and bytes distinct and acknowledge the sizing assumptions.
Of 100,000 membership checks, 80% are absent. With a 1% Bloom false-positive rate, how many exact preliminary reads remain?
Reveal a model answer
“Of 100,000 checks, 80,000 are absent. At a 1% false-positive rate about 800 absent checks still reach the database, alongside 20,000 present checks. That is about 20,800 reads instead of 100,000.”
Interviewer follow-up
Does it save the new URL’s durable insert too?
Reveal the follow-up answer
“No. It removes a preliminary read, while the exact claim or insert remains necessary.”
What the answer must demonstrate: Do not confuse lookup reduction with eliminating all authoritative work.
Bloom key A sets bits 2 and 7; B sets 7 and 12. Why can deleting A not simply clear its bits?
Reveal a model answer
“B shares bit 7, so clearing it can turn B into a false negative. Ordinary Bloom bits do not record ownership. I need a correctly managed counting variant or a rebuild/epoch policy.”
Interviewer follow-up
Can a counting filter delete any item that tests positive?
Reveal the follow-up answer
“No. A positive may itself be false, so decrementing for a never-inserted item can damage other entries. Deletions require reliable membership and accounting.”
What the answer must demonstrate: Deletion changes the guarantee unless ownership is accounted for.
“No. HyperLogLog estimates distinct cardinality; it cannot answer whether a particular URL was seen or enumerate URLs. It is useful for aggregate crawler statistics, while exact claim decisions need an exact set or database.”
Interviewer follow-up
Is 0.81% a maximum error at 16,384 registers?
Reveal the follow-up answer
“No. It is an approximate relative standard error from the classic analysis, not a deterministic per-answer bound.”
What the answer must demonstrate: Separate an aggregate estimator from a membership structure.
“Each counter contains the item’s own increments plus collisions. Under nonnegative insert-only updates, taking the minimum reduces collision inflation without dropping below the true count. In the example, min(27,23,22) estimates a true count of twenty as twenty-two.”
Interviewer follow-up
Can it list the busiest hosts by itself?
Reveal the follow-up answer
No. It answers estimates for supplied keys; heavy-key discovery needs candidate tracking. Also inspect the additive bound against total stream volume: epsilon = 0.001 at one million increments allows error of 1,000 for a fixed key, which can swamp a rare count.
What the answer must demonstrate: Qualify the update model and distinguish estimation from enumeration.
“Yes, when the sketch types, dimensions, hash functions and item normalization are compatible. Bloom union uses OR; HyperLogLog uses register maxima; Count-Min sums counters.”
Interviewer follow-up
What could still make the merged answer misleading?
Reveal the follow-up answer
“Different URL normalization or duplicate event delivery changes the represented data. HLL counts distinct identities while Count-Min counts occurrences, so their response to replay differs.”
What the answer must demonstrate: Compatible arrays are not enough; semantics must match.
What is an inverted index? If reset maps to {D2,D4} and access to {D1,D4}, how does reset AND access execute?
Reveal a model answer
An inverted index maps a term to the documents containing it. Here reset maps to D2 and D4, while access maps to D1 and D4. Intersecting the posting lists returns D4 without scanning every document body. Positions support phrases and term statistics support ranking.
Vector retrieval compares compatible numeric embeddings under a similarity measure and can match related wording without identical terms. It does not replace exact identifier fields or prove that a document is correct or authorized.
What the answer must demonstrate: Build the two lists and distinguish lexical matching from similarity.
“It helps retrieve semantically related wording, such as lost phone matching authenticator recovery. It is weaker for some precise identifiers and does not establish truth or permission, so I evaluate it alongside lexical search and metadata filters.”
Interviewer follow-up
Can I change embedding models without rebuilding vectors?
Reveal the follow-up answer
Only with an explicitly compatible representation contract. Otherwise stored and query vectors no longer share a meaningful space and need migration.
What the answer must demonstrate: Similarity is a retrieval signal.
Approximate search can miss neighbors that an exact result would include, in exchange for less work on suitable workloads. I compare it with an exact baseline under the same metric and eligibility filters, then tune latency, memory and recall together. Exactness does not require a full scan if an index can safely prove which candidates cannot win.
Interviewer follow-up
Does 99% neighbor recall imply 99% useful answers?
Reveal the follow-up answer
No. It measures approximation relative to the chosen vector metric, not whether the model or document collection captures user relevance.
What the answer must demonstrate: Separate approximation quality from semantic quality.
Estimate raw storage for ten million 768-dimensional float32 vectors.
Reveal a model answer
“Each vector is 768 × 4 = 3,072 bytes. Ten million require 30.72 GB in decimal units before graph links, metadata, text, and replicas. I would size those separately and benchmark any quantization loss.”
Interviewer follow-up
Does adding two replicas double or triple total copies?
Reveal the follow-up answer
Two additional replicas plus the original means three copies; clarify terminology before multiplying.
What the answer must demonstrate: Keep units and overhead explicit.
Why not add a keyword score directly to a cosine score?
Reveal a model answer
“Their scales and distributions differ. I can calibrate a learned combination or start with rank fusion, then evaluate. Reciprocal rank fusion uses positions in each result list and avoids pretending unlike raw scores have the same meaning.”
It spends more compute comparing the query with a smaller candidate set; it cannot recover a relevant document that never became a candidate unless another retrieval stage adds it.
What the answer must demonstrate: Candidate recall bounds reranking.
Why can filtering the final top twenty return no useful result?
Reveal a model answer
“All twenty may belong to another tenant even though relevant authorized documents exist deeper in the collection. I apply an eligible-document retrieval strategy and evaluate selective filters. In every case I enforce authorization before content leaves the trusted retrieval boundary.”
Interviewer follow-up
Can filtering only displayed citations secure an assistant?
Reveal the follow-up answer
No. Unauthorized snippets may already have entered its context and influenced the answer.
What the answer must demonstrate: Distinguish candidate starvation from data exposure.
Three of five returned documents are relevant; ten relevant documents exist. What are precision and recall?
Reveal a model answer
“Precision@5 is 3/5, or 60%. Recall@5 is 3/10, or 30%. I also measure ranking quality because users often inspect only the first results.”
Interviewer follow-up
What query set should the evaluation include?
Reveal the follow-up answer
Realistic exact IDs, paraphrases, rare cases, language variation, empty-result cases, and permissions—not just easy queries chosen to flatter the system.
What the answer must demonstrate: Use the correct denominator.
A document was deleted but remains searchable. How do you fix the contract?
Reveal a model answer
Send versioned deletion markers to the index and measure cleanup delay. To block access immediately, do not rely only on that delayed index update. Check current source permissions for the exact document and policy version before fetching its body or sending it to a model. Fetch that fixed content version; if versions differ, check permission again. Apply the promised access check when releasing the response too.
Authentication establishes a caller’s identity. Authorization checks a specific action on a specific resource. Tenant isolation requires those checks and data boundaries to prevent cross-tenant exposure through every path. For example, authenticated principal U10 belongs to Birch and must not read Acme invoice I17 through the API, cache, search, export, or file endpoint.
Interviewer follow-up
Would using unguessable invoice UUIDs remove the need for object authorization?
Reveal the follow-up answer
No. IDs can leak, be shared, or appear in logs. The service must check the actor, tenant, resource, and action regardless of how hard the ID is to guess.
What the answer must demonstrate: Use an actual permitted and forbidden resource path.
“It can treat it as a requested tenant, then verify the authenticated principal’s membership and permission. I never let a caller-selected tenant ID bypass that decision, and I propagate the verified scope into data access.”
Interviewer follow-up
What about a background worker?
Reveal the follow-up answer
It receives authenticated job context and an explicit execution/delivery authorization policy, not an unvalidated tenant string.
What the answer must demonstrate: Trace how the scope becomes trusted.
Two tenants may share invoice I17, and users may have different field permissions. I key cached bodies by tenant and immutable representation version, authorize the exact version and current policy scope, then return only that authorized representation. If the fetched body or required policy revision differs from the decision, I reauthorize or withhold it.
“It is useful defense in depth when policies, roles, and connection context are correct. I still enforce object/action permission in the application and verify privileged-role bypass behavior. A database policy cannot secure an unscoped object-storage or cache path.”
Interviewer follow-up
What if the app connects as the table owner?
Reveal the follow-up answer
In PostgreSQL owners normally bypass row security unless forced, and superusers/BYPASSRLS roles remain privileged. I use a restricted runtime role and transaction-local tenant context, then test reads and WITH CHECK behavior on writes through the actual pooled connections.
What the answer must demonstrate: Know the enforcement boundary.
“No. If the application can decrypt both tenants’ records, it can still send the wrong one. Check tenant permissions, route to the correct data and return only allowed fields. Encryption protects stored bytes; it does not make those application decisions.”
Interviewer follow-up
Where should keys live?
Reveal the follow-up answer
In protected key or secret infrastructure with scoped access and rotation, outside source control and client bundles.
What the answer must demonstrate: Name the threat each mechanism addresses.
Is a signed download URL private to the logged-in user?
Reveal a model answer
“Usually it is a bearer capability, so another person holding it can use it until its conditions expire. I authorize before issuance, limit scope and lifetime, and use an application-mediated access check when immediate revocation is required.”
Interviewer follow-up
Should URLs appear in ordinary logs?
Reveal the follow-up answer
Avoid recording capability tokens or query strings that expose access; use safe resource identifiers for observability.
What the answer must demonstrate: Possession can confer access.
How do you stop one tenant’s exports slowing every customer?
Reveal a model answer
“I bound per-tenant concurrency and total queues, schedule fairly, and separate heavy export workers from interactive reads. Quotas describe an enforceable budget; admission control prevents accepting more work than we can serve.”
Interviewer follow-up
What happens above the quota?
Reveal the follow-up answer
Return a clear retry or asynchronous scheduling contract rather than letting memory and latency grow without a bound.
What the answer must demonstrate: Security includes resource isolation.
“Use two tenants and attempt cross-tenant reads, writes, search, exports, attachment downloads, and cache hits. Also test revoked membership and expired capabilities. Each denied operation must leave data and side effects unchanged under its contract.”
Interviewer follow-up
Can error messages leak information?
Reveal the follow-up answer
Yes. Choose a consistent external disclosure policy while retaining detailed protected audit information for operators.
What the answer must demonstrate: Exercise alternate access paths.
RPO, recovery point objective, is the targeted maximum data-loss window. RTO, recovery time objective, is the targeted time to restore the agreed service. If the last recoverable state is 5 seconds before a disruption and service returns 7 minutes later, those are separate data-loss and restoration measurements to compare with the objectives.
Interviewer follow-up
Does observing 5 seconds of replica lag guarantee a 5-second RPO?
Reveal the follow-up answer
No. It is an observation under one condition. The design needs a survival and recovery mechanism that supports the target under its stated failure assumptions; lag may grow during a worse outage.
What the answer must demonstrate: Keep the two objectives separate and distinguish targets from measured guarantees.
Can asynchronous regional replication promise zero loss of acknowledged writes?
Reveal a model answer
“Not by itself. East can acknowledge a write and fail before it reaches West. To survive that regional loss without losing acknowledged writes, acknowledgement must require a durable copy or quorum outside East, within the stated failure model.”
Interviewer follow-up
What is the price?
Reveal the follow-up answer
Commit must wait for surviving remote durable state under a safe protocol, adding latency and possible refusal during partitions. Also inspect voting placement: two of three voters in one region can acknowledge a majority that disappears with that region.
What the answer must demonstrate: Place the acknowledgement boundary.
The East primary stops responding to West. Why is that alone insufficient to promote West safely?
Reveal a model answer
A failed health check cannot prove that East stopped writing. In the asynchronous two-region design, I require a verified stop or removal of its write capability before promotion; if that is impossible, writes remain paused. Alternatively, a proven quorum protocol prevents the isolated minority from committing. A new epoch stored only in West is not sufficient fencing.
“It can improve continuity for some operations, but shared dependencies and write conflicts remain. I would state whether each key has one home writer, uses global coordination, or permits a defined merge. Inventory cannot simply merge arbitrary decrements without a rule.”
“A snapshot alone cannot: it can leave almost a day of changes absent. Continuous recoverable logs or another replication mechanism may narrow that gap. I also need to test restore and replay time against the RTO.”
Interviewer follow-up
How long does transferring 6 TB at 1 GB/s take?
Reveal the follow-up answer
About 6,000 seconds, or 100 minutes, before other recovery work, using decimal units.
What the answer must demonstrate: Check both freshness and duration.
The standby has one fifth of peak capacity. Is failover ready?
Reveal a model answer
“Only if the recovery contract allows bounded degradation and the remaining capacity or scaling is verified. I would test databases and dependencies too, prioritize essential operations, and limit admission rather than overload the new primary.”
“Replicas can copy an accidental deletion or corruption. Backups and point-in-time recovery preserve earlier states under a separate protection policy. I would regularly restore and validate business records, not only check the backup job status.”
Interviewer follow-up
What if encryption keys are unavailable?
Reveal the follow-up answer
The bytes may be intact but unusable; key recovery is a dependency in the restore exercise.
What the answer must demonstrate:Replication is not historical recovery.
An old primary region recovers after failover. Why should writes not immediately be routed back?
Reveal a model answer
“West has accepted new writes, so keep it in charge. Bring East up to date or rebuild it from West, verify the data, then switch writers through a controlled handover. Prevent the former writer from continuing. If the regions have conflicting histories, resolve them before switching.”
Interviewer follow-up
How do you measure success?
Reveal the follow-up answer
Run a real read/write/reconciliation check and confirm the agreed capacity, data state, and SLO, rather than checking only that processes are up.
What the answer must demonstrate:Failback is a controlled state transition.
An SLI (service-level indicator) is a measured user-visible behavior. An SLO is its target over a window. An SLA is an agreement about service commitments and consequences, often contractual. An error budget is the amount of failure permitted by the SLO.
For a photo service, measure the fraction of eligible accepted photos that become ready within 60 seconds, after each photo's evaluation period has elapsed. Set an illustrative SLO of at least 99.9% over 30 days. Of one million evaluated photos, at most 1,000 can miss the target. A separate SLA could specify a contractual commitment and credits; it need not use the same threshold as the internal SLO. Also measure acceptance so the service cannot make its completion ratio look good by rejecting every upload.
Interviewer follow-up
Why wait for the evaluation period to elapse?
Reveal the follow-up answer
“A photo accepted five seconds ago has not yet missed a sixty-second deadline. I evaluate it once when that deadline passes and count the durable ready-by-deadline result. The rolling window uses those evaluation times; unfinished jobs and missing telemetry must not disappear from the denominator.”
What the answer must demonstrate: A percentage without a denominator and window is incomplete.
Could accepting no uploads make your completion SLO look perfect?
Reveal a model answer
“Yes, if it treats no data as success, or shows only accepted uploads and hides rejected attempts. Zero evaluated photos proves nothing about completion. Also measure how many valid attempts are accepted, signal missing data and use a test upload when useful. Then we can distinguish failure to accept uploads from failure to process them.”
Interviewer follow-up
Should malformed uploads count as service failures?
Reveal the follow-up answer
“That depends on the specified contract, but I would separate expected validation rejection from failures of valid requests and avoid exclusions that hide our defects.”
What the answer must demonstrate: Beware metrics that improve by refusing useful work.
For 1,000,000 evaluated operations and a 99.9% success SLO, what is the error allowance? What burn rate does a 2% bad-event rate represent?
Reveal a model answer
“At 99.9%, one million evaluated operations allow one thousand bad events. If a recent window has two percent bad events against a 0.1 percent allowance, its burn rate is twenty. I would interpret that with traffic and window size before deciding how urgently to page.”
Interviewer follow-up
Why use more than one alert window?
Reveal the follow-up answer
“A short window detects rapid deterioration; a longer one helps establish that it persists. The combination reduces both slow detection and noisy reaction to tiny samples.”
What the answer must demonstrate: Keep percentage points and ratios distinct.
A photo takes 100 seconds: 95 queued, 4 rendering, 1 publishing. What does this trace reveal that low CPU usage does not?
Reveal a model answer
“The trace assigns ninety-five seconds to waiting, four to rendering, and one to publication. Low CPU cannot tell me whether concurrency is accidentally restricted or demand is absent. The trace locates the delay; release/configuration evidence then helps identify the cause.”
Interviewer follow-up
Does a high queue age prove the broker is faulty?
Reveal the follow-up answer
“No. Slow or insufficient workers can produce the same symptom. I would inspect service rates and stage behavior rather than blame the queue by its name.”
What the answer must demonstrate: Separate symptom, location, and causal evidence.
“Metrics show aggregate trends and support alerts. Structured logs record individual events. Traces or correlated job events connect a specific journey across stages. For P501 I use metrics to detect late completion and the trace plus targeted logs to explain where it waited and which release handled it.”
Interviewer follow-up
Should photo IDs be labels on every metric?
Reveal the follow-up answer
“Usually not. That creates unbounded time-series cardinality. Keep per-photo details in appropriately protected logs or traces.”
What the answer must demonstrate: Choose the evidence type according to the question.
What should the canary compare before full deployment?
Reveal a model answer
“Compare similar workloads, user-visible completion, throughput and stage delay, with enough observations to distinguish a signal from noise. For worker changes, separate canary and control pools can make queue-wait attribution meaningful. If both versions share a queue, rising age affects the cohort comparison and cannot by itself blame one version. I would inspect per-version throughput/configuration and stop expansion or revert when the canary violates the agreed guardrails.”
Interviewer follow-up
Can every deployment be rolled back by restoring old binaries?
Reveal the follow-up answer
“No. Incompatible data/schema changes may make old code unsafe. I would plan compatible transitions and a recovery path before rollout.”
What the answer must demonstrate: Deployment safety includes data compatibility.
The old worker version is back. Can you close the incident?
Reveal a model answer
“Only after verifying the backlog drains and timely completion recovers. I also check that retrying a job did not publish duplicate results or repeat other side effects, and that photo permissions remain correct. Restoring the old version is a recovery step; I still need to verify that users can upload and view photos successfully.”
Interviewer follow-up
What if arrival rate still equals processing capacity?
Reveal the follow-up answer
“Existing backlog will not drain. I need temporary spare capacity or reduced admission and must communicate the ongoing delay.”
What the answer must demonstrate: Verify recovery under continuing load.
“RPO tells me how much recent accepted work may be lost; RTO tells me how soon useful service should return. I would measure both during a drill and verify original files, records of pending jobs, publication status, and permissions, rather than timing only a database restore command.”
“No. Replicas can carry the same mistake, and recovery has dependencies beyond copying state. We need to verify that uploads, background processing, and authorized viewing all work afterward.”
What the answer must demonstrate: Recovery objectives apply to the service outcome.
Two requests can still generate the same candidate. The database’s unique constraint accepts only one competing insert; the loser chooses another code.
Interviewer follow-up
Why not check first?
Reveal the follow-up answer
Both requests can observe absence before either inserts, so a check followed by an unguarded write races.
What the answer must demonstrate: Names an atomic uniqueness constraint and the check-then-write race.
Scope a request key to the authenticated owner and store its payload identity and result atomically with the mapping. Replay a matching completed request.
Interviewer follow-up
What if the payload differs?
Reveal the follow-up answer
Reject the reused identity as a conflict; it represents a different operation.
What the answer must demonstrate: Keeps request identity, payload validation and mapping commit connected.
Why is an old upload not automatically safe to delete?
Reveal a model answer
Its uploader may still be completing. Cleanup must first mark the attempt canceled in the same database state that publication checks, then delete its bytes.
Interviewer follow-up
What does a late uploader do?
Reveal the follow-up answer
Its READY transition fails once the attempt is canceled; it cannot bypass that guard.
What the answer must demonstrate: Identifies a state guard shared by publication and deletion.
The server verifies that chunks exist and cannot be deleted, checks current permission and the expected base revision, then commits the new revision in metadata.
Interviewer follow-up
What if the response is lost?
Reveal the follow-up answer
Retry the same operation identity and retrieve its committed result.
What the answer must demonstrate: Separates byte upload from publication and request replay.
A unique user/post pair makes repeated like and unlike requests well-defined. Blind counter increments duplicate actions after retries.
Interviewer follow-up
Must the displayed count update synchronously?
Reveal the follow-up answer
Not for this scope; the displayed count may catch up as relationship changes are processed, while the saved user/post row tells us whether that user has liked the post.
What the answer must demonstrate: Distinguishes authoritative relationship from derived count.
First-frame delay, rebuffered fraction, device/codec errors and actual decoded segment checks reveal user experience better than metadata response codes.
Interviewer follow-up
Which cost tradeoff follows?
Reveal the follow-up answer
Additional encoding may reduce watched bytes, but only if quality and device compatibility remain acceptable.
What the answer must demonstrate: Measures decoded user experience and explicit cost tradeoffs.
Follow character edges to the prefix node, then use descendant terminal terms. A terminal marker distinguishes a complete term from an intermediate path.
Interviewer follow-up
Why is cap both a result and an internal node?
Reveal the follow-up answer
It is a complete term and a prefix of capital, captain and caption.
What the answer must demonstrate: Explains prefix traversal and terminal semantics.
Build the replacement separately so a query never combines a changed term score with an old prefix shortlist. Each request uses one complete version of the terms, normalization rules and scores.
Interviewer follow-up
What memory trap appears during rollout?
Reveal the follow-up answer
Both the current and replacement indexes may need to coexist, plus process reserve.
What the answer must demonstrate: States coherent data and simultaneous-version memory costs.
What must be clarified before choosing a limiter algorithm?
Reveal a model answer
Ask whose allowance is shared, which action consumes it, over what interval, whether bursts are allowed and whether an outage should admit or deny. Those answers determine the algorithm and stored state.
Interviewer follow-up
Does a failed conversion refund allowance here?
Reveal the follow-up answer
No. This contract counts accepted admissions, not successful conversions.
What the answer must demonstrate: States all dimensions of the product contract before storage.
A promoted replica may lack an acknowledged admission and grant extra allowance. The selected datastore must preserve committed usage or stop admission.
An indexed match now has different text. What should be returned?
Reveal a model answer
Authorize and load the same current version, then verify it still satisfies the query. Omit a nonmatch; the index may temporarily miss new matches too.
Interviewer follow-up
Can a cached public permission authorize a newly private version?
Reveal the follow-up answer
No. The permission decision must bind the returned content version.
What the answer must demonstrate: Binds content and permission versions and acknowledges recall lag.
Why do per-worker delays fail to enforce politeness?
Reveal a model answer
Multiple workers can each obey their own delay while sending concurrent requests to the same origin. Origin-wide state must coordinate starts and active requests.
No. A timeout can leave the request outcome unknown, and recovery may repeat the network operation. Durable idempotent processing is the defensible promise.
Interviewer follow-up
What prevents late completion from overwriting a newer attempt?
Reveal the follow-up answer
A conditional state update checks the current attempt token.
What the answer must demonstrate: Allows duplicate network attempts while guarding state updates.
Ordinary active audiences benefit from prepared references, while celebrity fanout may create millions of unused writes. Pull those histories during actual reads.
Interviewer follow-up
What decides the threshold?
Reveal the follow-up answer
Posting rate, active audience reads and measured write/merge costs, not follower count alone.
What the answer must demonstrate: Uses audience economics to justify both paths.
Distribution interest, approved friendship and private-group membership grant different behavior and access. A one-way follow cannot silently authorize friends-only content.
Interviewer follow-up
Does unfollowing make a public profile private?
Reveal the follow-up answer
No. It removes that source from this followed-content feed; profile access follows its own policy.
What the answer must demonstrate: Distinguishes distribution from actual access grants.
Why can a café across a cell boundary be the closest result?
Reveal a model answer
Cells organize storage; they do not constrain physical distance. Maya at x=990 and P12 at x=1010 are twenty meters apart. Search must cover intersecting cells and apply exact distance.
Interviewer follow-up
Can a bounding rectangle be the final radius answer?
Reveal the follow-up answer
No. Its corners lie outside the circle. Use it to find candidates, then enforce the requested distance.
What the answer must demonstrate: Explain cross-cell coverage and the exact circular distance predicate.
After examining all eligible bounded candidates, or after every remaining region has a valid minimum distance worse than the current twentieth result. Equal distances need the stated tie-breaker.
Interviewer follow-up
Why filter closed places before pruning?
Reveal the follow-up answer
A closed café is not an eligible result and cannot justify excluding a farther open café.
What the answer must demonstrate: State the stopping bound and apply eligibility before counting best results.
It makes the distance unit explicit. Latitude and longitude are angles, and longitude distance varies with latitude, so raw degree subtraction is unsuitable.
Interviewer follow-up
Which implementation would you start with?
Reveal the follow-up answer
A mature geography-aware spatial database and indexed distance predicate, with boundary tests.
What the answer must demonstrate: Distinguish angular coordinates from distance units and use appropriate geography.
Can current detail reads repair every stale-index error?
Reveal a model answer
They can remove incorrect returned candidates but cannot discover a moved place absent from the new region. Preserve the indexing-delay contract or use a caught-up index.
Interviewer follow-up
Why not substitute current coordinates during pruning?
Reveal the follow-up answer
Old region bounds may no longer contain the moved point, invalidating the stopping argument.
What the answer must demonstrate: Distinguish rejecting stale candidates from discovering missing moved records.
Fetch the viewer’s authorized contacts, read their latest positions in a batch, check expiry and compute distances. A bounded contact set may not need a global spatial index.
Interviewer follow-up
What happens after sharing is revoked?
Reveal the follow-up answer
Subsequent serving checks must deny access even if old coordinates or candidate entries remain cached.
What the answer must demonstrate: Use current viewer consent and presence expiry, including the small-contact-set alternative.
Position describes an observation; availability is a durable promise about whether the driver can accept work. A fresh heartbeat must not release an active assignment.
Interviewer follow-up
Can the map be stale?
Reveal the follow-up answer
Yes under an explicit freshness policy, because acceptance rechecks current authority.
What the answer must demonstrate: Separate observed location from authoritative assignment availability.
D17 accepts R501 and R502 concurrently. What prevents two riders winning?
Reveal a model answer
Both transactions lock or conditionally guard the same driver record and atomically update the corresponding ride. The second sees D17 assigned and changes neither side.
Interviewer follow-up
What if two different drivers accept R501?
Reveal the follow-up answer
Both must also guard the ride row; checking only driver availability misses that race.
What the answer must demonstrate: Protect both driver exclusivity and ride exclusivity in the same transaction.
The acceptance reply is lost. What should the retry return?
Reveal a model answer
The saved assignment for the same authenticated operation or accepted offer. Check that result after lock acquisition so a concurrent duplicate observes the committed winner.
Interviewer follow-up
Why not return driver unavailable?
Reveal the follow-up answer
The caller may be retrying their own successful assignment, not proposing a second ride.
What the answer must demonstrate: Recover the saved assignment after lock acquisition instead of rejecting a successful retry.
Why include a session generation as well as a sequence?
Reveal a model answer
Sequences order updates within one publishing session. A generation distinguishes a restarted or replacement session and fences delayed packets from the old one.
Interviewer follow-up
Does receipt time prove GPS freshness?
Reveal the follow-up answer
No. It measures arrival; the observation time and GPS accuracy have separate validation limits.
What the answer must demonstrate: Explain session replacement, sequence ordering and observation-time limits.
Both use the ride authority. Cancellation first blocks acceptance; assignment first means cancellation follows the assigned-trip policy and releases both references atomically if allowed.
Interviewer follow-up
Can notification order decide the outcome?
Reveal the follow-up answer
No. Clients use committed state and versions, because notifications can arrive late or twice.
What the answer must demonstrate: Serialize cancellation with acceptance and use committed versions for notifications.
Why not change assignment owner on every spatial-cell crossing?
Reveal a model answer
Cells help find nearby drivers. Assignment ownership keeps ride and driver records consistent. Moving those records on every cell crossing would add unnecessary coordination.
Interviewer follow-up
What changes for cross-region matching?
Reveal the follow-up answer
Introduce an explicit reservation or transfer protocol, or a distributed transaction, before claiming the same exclusivity.
What the answer must demonstrate: Distinguish spatial movement from regional assignment-authority transfer.
A hold is a durable reservation with a business deadline. A row lock protects a short state transition; holding it while a person pays wastes connections and does not define recovery.
Interviewer follow-up
What if cleanup stops?
Reveal the follow-up answer
The deadline still makes the hold invalid; cleanup only restores availability promptly.
What the answer must demonstrate: Distinguish minutes-long business reservations from short database locks.
How do seats 54–56 and 56–57 avoid a partial allocation?
Reveal a model answer
Each transaction locks its selected seats in order, validates the full set and commits all changes together. The loser on seat 56 rolls back its entire set.
Interviewer follow-up
Would atomic updates on individual seats suffice?
Reveal the follow-up answer
No. They might leave a customer holding only some requested seats.
What the answer must demonstrate: Validate and commit the entire requested seat set or roll it back.
Does a waiting room guarantee first-arrival seat allocation?
Reveal a model answer
No. Several admitted clients can race and network timing can reorder their commits. State whether the product promises bounded load or strict fairness.
Interviewer follow-up
What does strict fairness cost?
Reveal the follow-up answer
Durable ordered grants and less concurrency, with possible head-of-line blocking.
What the answer must demonstrate: State the chosen admission fairness contract and its concurrency cost.
How does the design change when the booking is a three-night hotel stay?
Reveal a model answer
Inventory becomes a room-type capacity row for every occupied hotel-local date in [checkIn, checkOut). One transaction checks and reserves all nights. If the middle night is unavailable, reject the entire stay with no partial holds. Confirmation or release changes every nightly allocation and the reservation state atomically.
Interviewer follow-up
Does booking each night in a separate successful transaction preserve this contract?
Reveal the follow-up answer
No. It can commit the first and third nights while the second fails, leaving a partial stay. Keep one hotel’s night rows in the same transaction domain, or explicitly choose a more complex pending reservation workflow across independent owners.
What the answer must demonstrate: Identify room-type/date inventory, exclusive checkout date, whole-stay atomicity and duplicate-safe hold transitions.
The version identifies state to replace; the request ID identifies one attempted change. Their combination supports conflict detection and retry recovery.
Interviewer follow-up
Can one request ID carry different values?
Reveal the follow-up answer
Reject that reuse by comparing the saved payload fingerprint.
What the answer must demonstrate: Distinguish expected state from attempted operation and reject conflicting ID reuse.
Two clients replace version 7. Why can only one win?
Reveal a model answer
The agreed command order applies one check-and-update first, advancing the version. The next command checks the new current state and records conflict.
Interviewer follow-up
What is the unsafe alternative?
Reveal the follow-up answer
Comparing outside the atomic mutation path lets both callers pass the old check.
What the answer must demonstrate: Place version comparison and mutation inside the agreed serialized application step.
Why might a process labeled leader be unable to serve a strong GET?
Reveal a model answer
It may be isolated while another group has elected a successor. The leader must confirm it still leads and apply the committed commands required for the read.
Interviewer follow-up
Can it offer an old value anyway?
Reveal the follow-up answer
Only through an explicitly weaker stale-read contract.
What the answer must demonstrate: Require current read authority and sufficient applied state after leadership change.
One business event can create several channel deliveries, and each delivery may require several calls. Separate identities preserve partial status and safe retries.
Interviewer follow-up
Which identity should a provider key normally represent?
Reveal the follow-up answer
The logical delivery, not each transport attempt.
What the answer must demonstrate: Keep business intent, channel delivery and transport attempt identities separate.
The email call times out after possible acceptance. How do you retry?
Reveal a model answer
Recover the same delivery through supported idempotency or lookup. Without either, keep unknown and apply the category’s explicit duplicate-risk policy.
Interviewer follow-up
Can switching providers fix it?
Reveal the follow-up answer
No. A second provider cannot deduplicate an effect at the first.
What the answer must demonstrate: Use original provider identity or lookup and state the unsupported-provider limit.
How does the ledger differ from payment workflow state?
Reveal a model answer
Workflow records which actions are pending or known complete. Immutable journals record financial movements with balanced debit and credit totals per currency.
Interviewer follow-up
Can a posted journal be edited to correct it?
Reveal the follow-up answer
Use an auditable reversing or correcting journal rather than erasing history.
What the answer must demonstrate: Distinguish mutable workflow state from immutable balanced accounting journals.
The processor captured but the worker lost its reply. What happens?
Reveal a model answer
Keep the stored operation unknown, recover through its provider identity or lookup, and apply verified evidence locally. Do not issue a new charge to repair missing local status.
Interviewer follow-up
What if provider key retention expired?
Reveal the follow-up answer
Reconcile or review the old operation; an old local key cannot extend the processor’s deduplication window.
What the answer must demonstrate: Recover the original processor effect without issuing another charge to repair local state.
What must change when the prompt asks Alice to transfer existing wallet funds to Bob?
Reveal a model answer
Recover an authorized matching retry by its stable transfer ID before applying new-transfer eligibility checks. For a new positive same-currency transfer, authenticate Alice’s debit authority, lock both accounts in stable order, recheck spendable funds after outgoing holds, and commit the balanced journal, both balances and saved result together. This internal transfer does not need a card-processor call.
Interviewer follow-up
Why do balanced journal entries alone not prevent overdraft?
Reveal the follow-up answer
Two transfers can each balance their debit and credit while both spending the same stale source balance. Serialize the current spendable-funds check with each debit. If the accounts are on different shards, use an actual distributed transaction or an explicitly pending, reserved-funds workflow.
What the answer must demonstrate: Separate internal transfers from merchant capture; protect spendable balance and journal atomicity under concurrent spends and retries.
Two users can replace the same base with different complete copies, causing the later save to erase independent work. Operations preserve the intent the merge algorithm needs.
How do X and Y inserted at position 1 both survive?
Reveal a model answer
A defined tie-break accepts one first and transforms the other position around it. Both clients reconcile pending and accepted edits under the same rules.
Interviewer follow-up
Does the example implement a full editor?
Reveal the follow-up answer
No. Deletes, undo, Unicode and overlapping operations require a proven complete algorithm.
What the answer must demonstrate: Trace both client reconciliation and deterministic server transformation without generalizing an insert demo.
How can an edit disappear between opening and subscribing?
Reveal a model answer
An operation may commit after the initial snapshot/read but before live subscription. Replay or buffer from the last applied version to cover that gap.
Interviewer follow-up
What if version 24 arrives before 23?
Reveal the follow-up answer
Fetch the missing history before applying later positional edits.
What the answer must demonstrate: Close the snapshot-to-subscription gap and repair missing versions before positional application.
Why must storage reject a stale document coordinator?
Reveal a model answer
The old process may resume with open sockets after a replacement takes over. Checking its ownership generation at append prevents a second accepted history.
Interviewer follow-up
Why check the head too?
Reveal the follow-up answer
A transform computed against old content must be recomputed after intervening operations.
What the answer must demonstrate: Check ownership generation and expected head at the protected append boundary.
Can a snapshot replace all operation history immediately?
Reveal a model answer
It can reconstruct current text but may not preserve the context needed to transform supported old pending edits. Retention follows the reconnect contract.
Interviewer follow-up
What happens beyond that contract?
Reveal the follow-up answer
Preserve the local draft and require an explicit current-snapshot merge workflow.
What the answer must demonstrate: Retain required transform context or explicitly preserve drafts for resynchronization.
Metrics summarize aggregate behavior, logs describe events and traces link timed work across a request. Each has different identity, storage and query needs.
Interviewer follow-up
How do they connect?
Reveal the follow-up answer
Use consistent service context and request/trace IDs in the detailed evidence.
What the answer must demonstrate: Connect aggregate symptoms, individual events and request-span relationships.
Each original series can reset independently. Reset-aware rates preserve those boundaries; summing raw counters first can hide resets and create false activity.
Interviewer follow-up
Does a counter value equal new events in the last minute?
Reveal the follow-up answer
No. It is cumulative state whose change over an interval must be interpreted.
What the answer must demonstrate: Handle each counter reset before combining rates across instances.
The schedule is the rule, an occurrence is one intended resolved run, and attempts retry that same run. This keeps history and business identity stable.
Interviewer follow-up
Can a retry use current edited parameters?
Reveal the follow-up answer
No. An already created occurrence retains its frozen revision.
What the answer must demonstrate: Preserve one intended occurrence and its frozen parameters across attempts.
It requests cooperative stopping and prevents later accepted completion/retries once cancellation becomes terminal. External effects already started may still complete.
Interviewer follow-up
What if completion already committed?
Reveal the follow-up answer
Return the committed outcome rather than pretend cancellation erased it.
What the answer must demonstrate: Serialize terminal cancellation with completion and disclose already-started external effects.
Does broker transactional processing make an arbitrary payment exactly once?
Reveal a model answer
No. Its guarantee covers only the storage and outputs that participate in its documented transaction protocol; a payment provider needs its own stable identity and uncertain-outcome recovery.
Two checkouts each request two of three mugs. What prevents overselling?
Reveal a model answer
A short transaction locks or conditionally updates the same stock authority and checks the full line set. After one reserves two, the other sees only one available.
Capture may have succeeded but timed out. Should allocations expire?
Reveal a model answer
No. Keep the appropriate allocated claim, block shipping and reconcile the original payment attempt. A temporary hold deadline is not a release policy for possibly paid inventory.
Interviewer follow-up
What if capture definitively fails?
Reveal the follow-up answer
Cancel and release through guarded transitions, recording the resolved payment outcome.
What the answer must demonstrate: Keep allocations during unknown capture and reconcile the same financial attempt.
The order transaction commits fulfillment authorization and a unique shipment intent only if allocations and capture are valid. Cancellation must win before that boundary or use an explicit warehouse cancellation/return flow.
Interviewer follow-up
Why is a separate eligibility read insufficient?
Reveal the follow-up answer
Cancellation could release inventory between the read and an unchecked dispatch call.
What the answer must demonstrate: Order cancellation and fulfillment authorization in one transaction, and create one unique shipment intent.
What changes when lines belong to independent warehouses?
Reveal a model answer
There is no longer one all-line transaction. Persist local step outcomes and compensation while exposing pending state; confirm only after every required allocation exists.
Interviewer follow-up
Can rollback of the coordinator undo a remote hold?
Reveal the follow-up answer
No. Releasing it is a separate idempotent action at the inventory owner.
What the answer must demonstrate: State the loss of global atomicity and persist compensating local steps.
Why is event-ID deduplication insufficient for match awards?
Reveal a model answer
The same match may arrive under a different transport event ID. Store its canonical contribution and source revision so business meaning is applied once.
Interviewer follow-up
What commits together?
Reveal the follow-up answer
Accepted evidence, the contribution change, player total/version and outgoing projection work.
What the answer must demonstrate: Protect canonical match contribution/revision as well as transport-event identity.
Does waiting one second after cutoff prove an award board is complete?
Reveal a model answer
No. Use the declared late-result policy to establish which accepted inputs count or whether sources have sent all eligible results. Apply those inputs and verify the board before freezing awards.
Interviewer follow-up
What happens to a later fraud correction?
Reveal the follow-up answer
Publish a separate audited adjudication version rather than silently altering the historical award.
What the answer must demonstrate: Establish cutoff completeness and retain immutable award/adjudication evidence.
What is the first distinction in a maps interview?
Reveal a model answer
Separate drawing a map from computing directions. Tiles are reusable visual objects, while a route is a legal weighted path for particular endpoints and preferences. That distinction explains why CDN bandwidth and search CPU need different capacity estimates. Define departure time and access rules before choosing an engine.
Interviewer follow-up
Does a nearby road make a valid snap?
Reveal the follow-up answer
No. Direction, vehicle access, turns and actual connections matter; an overpass can be geographically close but inaccessible.
What the answer must demonstrate: Separate drawing a map from computing directions.
Its full cost is 3 + 8 = 11, while A-B-D costs 4 + 4 = 8. Dijkstra keeps alternative tentative distances, settles C and then B, and improves D before settling it. It does not commit to the cheapest first edge as an entire route. The proof assumes fixed nonnegative costs and a correct legal-state model.
Interviewer follow-up
Can changing live weights be read during this search?
Reveal the follow-up answer
Not under that static proof. Pin a compatible model or use a separately justified time-dependent algorithm.
What the answer must demonstrate: Its full cost is 3 + 8 = 11, while A-B-D costs 4 + 4 = 8.
It requires 100 CPU-seconds per second, or 100 fully busy cores before headroom. At 60% utilization the rough planning value is 167 cores before redundancy. This is not a deployment guarantee: long routes, memory locality and traffic mix must be benchmarked. Tile delivery remains a separate high-byte workload.
Interviewer follow-up
Why not add HTTP threads first?
Reveal the follow-up answer
Threads do not create CPU capacity or reduce graph exploration. Determine whether CPU, memory or network is limiting.
What the answer must demonstrate: It requires 100 CPU-seconds per second, or 100 fully busy cores before headroom.
How can an update avoid corrupting an active route?
Reveal a model answer
Build and verify immutable compatible artifacts, then activate their manifest atomically. A request acquires and retains one bundle reference. Cleanup cannot reclaim that bundle until its readers finish. This prevents an active search from combining old shortcuts with incompatible new weights, while permitting later requests to use the new version.
Interviewer follow-up
Is the manifest itself proof of compatibility?
Reveal the follow-up answer
No. Build validation and serving checks establish compatibility; the manifest records the set they validated.
What the answer must demonstrate: Build and verify immutable compatible artifacts, then activate their manifest atomically.
B-D closes while the eight-minute route is running. What happens?
Reveal a model answer
The worker validates the expanded road sequence against the closure version at its final boundary. If that version forbids B-D, it recomputes using a compatible model or returns an explicit retry. It cannot release a known-invalid route merely because the original graph was pinned. An incident reported afterward may require a later navigation refresh.
Why is choosing the nearest regional border unsafe?
Reveal a model answer
The minimum-cost legal route may cross a farther border, leave and reenter a region, or use a road that looks geometrically indirect. An overlay must represent valid interregional path costs and preserve the search problem. Regional boundaries are deployment boundaries, not road restrictions. Keep full graph replicas when their memory cost is acceptable.
Interviewer follow-up
What is the added operational cost?
Reveal the follow-up answer
Compatible overlay/local releases, cross-region coordination, more complex expansion and additional failure behavior.
What the answer must demonstrate: The minimum-cost legal route may cross a farther border, leave and reenter a region, or use a road that looks geometrically indirect.
First measure whether search CPU is the bottleneck. The chosen final design can scale independent local Dijkstra searches with warmed replicas. A* or precomputed shortcuts are further optimizations, tested against that reference. Explain their additional correctness and update assumptions only if the interviewer asks to extend the design.
Interviewer follow-up
What if updates rebuild too slowly?
Reveal the follow-up answer
Retain a correct fallback or choose an update-friendly method; faster queries do not compensate for failing the freshness contract.
What the answer must demonstrate: First measure whether search CPU is the bottleneck.
Use historical travel weights only if the product permits it, label freshness and preserve known access restrictions. Invalid topology or no legal connection requires failure or a verified base-search fallback, not straight-line directions. Monitor traffic age separately from query success so an available endpoint cannot conceal obsolete estimates.
Interviewer follow-up
What do you test before release?
Reveal the follow-up answer
Turns, one-way roads, overpasses, disconnected endpoints, border reentry and closure invalidation, plus representative long-route load.
What the answer must demonstrate: Use historical travel weights only if the product permits it, label freshness and preserve known access restrictions.
Validated identified clicks by occurrence time. Received requests, unique people and billable interactions have different rules, so the policy version and trusted provenance are part of the record. Transport deduplication does not establish that a distinct event is legitimate or billable.
Interviewer follow-up
Why keep raw input?
Reveal the follow-up answer
It explains revisions and supports bounded investigation or a controlled recount.
What the answer must demonstrate: Validated identified clicks by occurrence time. Received requests, unique people and billable interactions have different rules, so the policy version and trusted provenance are part of the record.
A collector durably appends identified events; one worker transactionally inserts the processed identity and increments the correct ad/minute counter. Queries read committed counts. Saving both effects together makes a lost worker response recoverable without assuming every delivery occurs once.
Interviewer follow-up
Why not expose a stateless increment endpoint?
Reveal the follow-up answer
A timed-out caller can repeat the increment even when the first one committed.
What the answer must demonstrate: A collector durably appends identified events; one worker transactionally inserts the processed identity and increments the correct ad/minute counter.
Where does C901 belong when it arrives at 09:02:05?
Reveal a model answer
Its occurrence at 09:00:58 assigns it to the 09:00 minute. If it is within the configured late-update policy, revise that original window from 99 to 100. Otherwise retain it for the correction process. Processing time must not silently answer a different business question.
Interviewer follow-up
What does a watermark mean?
Reveal the follow-up answer
Declared progress through event time, used to close windows under explicit assumptions; it does not rule out all late events.
What the answer must demonstrate: Its occurrence at 09:00:58 assigns it to the 09:00 minute.
The sink commits and the processor crashes before progress is saved. What happens?
Reveal a model answer
Recovery replays the event. The sink transaction finds its saved scoped identity and makes no second count change. If the previous transaction had aborted, neither identity nor effect would exist and replay would apply it. This contract must be verified for the actual external store.
Interviewer follow-up
Does a framework checkpoint automatically cover every database?
Reveal the follow-up answer
No. Use a documented transactional/checkpoint integration or a deliberately idempotent sink.
What the answer must demonstrate: Recovery replays the event. The sink transaction finds its saved scoped identity and makes no second count change.
First measure batching and database limits. If one counter remains hot, use stable partial keys, such as 16 salts averaging 2,500/s, then combine all partials before showing that ad’s total. Stable event identity still prevents duplicate contribution. Splitting every ordinary key adds unnecessary report and state cost.
Interviewer follow-up
What if a partial is unavailable?
Reveal the follow-up answer
Disclose incompleteness, use a labeled completed report or fail; do not treat missing data as zero.
What the answer must demonstrate: First measure batching and database limits. If one counter remains hot, use stable partial keys, such as 16 salts averaging 2,500/s, then combine all partials before showing that ad’s total.
Why can worker-local top lists miss an hourly winner?
Reveal a model answer
Red = 6 on two workers totals 12, but each worker can select a different local item at 7. A report must aggregate each ad’s complete total before sorting winners. This design uses a periodic report with stated build time; exact globally simultaneous rankings are a stronger follow-up.
Interviewer follow-up
Does a cursor create a fixed report?
Reveal the follow-up answer
No. Stable pagination needs a retained report or snapshot, not just a remembered key.
What the answer must demonstrate: Red = 6 on two workers totals 12, but each worker can select a different local item at 7.
How long does a ten-minute peak backlog take to recover?
Reveal a model answer
There are 120 million events. If processing resumes at 300,000/s while 200,000/s still arrive, spare throughput is 100,000/s and recovery takes 1,200 seconds. Preserve admitted work, expose lag and tighten intake before retention is exhausted.
Interviewer follow-up
Can progress be advanced to make the dashboard look healthy?
Reveal the follow-up answer
No. That can falsely finalize incomplete windows.
What the answer must demonstrate: There are 120 million events. If processing resumes at 300,000/s while 200,000/s still arrive, spare throughput is 100,000/s and recovery takes 1,200 seconds.
What happens to a very old resend after identity expiry?
Reveal a model answer
A24-hour online identity window cannot promise that an arbitrary historical resend is duplicate-safe. Enforce an admission policy for older work and use controlled reconstruction from retained evidence for historical repair. Do not silently treat forgotten identities as new clicks. Internal recovery obeys the same limit: stop automatic replay beyond retained identity coverage and rebuild affected counters from raw evidence instead of applying old inputs to existing counts.
Interviewer follow-up
When would you add exact global correction publication?
Reveal the follow-up answer
When reproducible audit requirements justify the additional snapshot, ownership and publication protocol.
What the answer must demonstrate: A24-hour online identity window cannot promise that an arbitrary historical resend is duplicate-safe.
Which identifier properties must be clarified first?
Reveal a model answer
Clarify namespace, integer width, allowed gaps, uniqueness and ordering. This design requires distinct integers but permits gaps and non-global issue order, so durable numeric ranges suffice. Record creation, authorization and safe business retries remain separate application responsibilities.
Interviewer follow-up
Does an allocated ID imply a committed order?
Reveal the follow-up answer
No. The order transaction can still fail or be abandoned.
What the answer must demonstrate: Clarify namespace, integer width, allowed gaps, uniqueness and ordering.
How does the single-row baseline avoid duplicates?
Reveal a model answer
A short transaction locks the namespace counter, reserves its next value, advances the counter and commits before returning. Concurrent callers serialize at that row. The durability/failover policy must retain acknowledged advances; unused values may become gaps.
Interviewer follow-up
What if the reply disappears after commit?
Reveal the follow-up answer
The value remains allocated and cannot be recycled merely because its use is uncertain.
What the answer must demonstrate: A short transaction locks the namespace counter, reserves its next value, advances the counter and commits before returning.
At ten million IDs/s it reduces authority operations to roughly 1,000 range reservations/s. Each process serves calls with a local synchronized cursor until its range ends. The tradeoff is abandoned space after crashes and loss of global issue order across independent ranges.
Interviewer follow-up
What would you benchmark?
Reveal the follow-up answer
Authority commit latency and failover plus local cursor contention and replenishment tails.
What the answer must demonstrate: At ten million IDs/s it reduces authority operations to roughly 1,000 range reservations/s.
Two threads request an ID simultaneously. Why is a range not enough?
Reveal a model answer
Disjoint ranges protect different processes, but both threads could read the same local cursor before either advances it. Use an atomic reservation or lock that checks the end and advances before returning. Batch calls obey the same boundary.
Interviewer follow-up
What happens at the exclusive end?
Reveal the follow-up answer
Switch to another committed range, wait within a deadline or fail; never issue outside the allocation.
What the answer must demonstrate: Disjoint ranges protect different processes, but both threads could read the same local cursor before either advances it.
How do you recover a generator without persisting every increment?
Reveal a model answer
Discard its entire old remainder and request a fresh disjoint range under a new incarnation. This allows gaps but avoids guessing which values were returned before the crash. A paused old process still holds a different range from its replacement.
Interviewer follow-up
What about cloning live memory?
Reveal the follow-up answer
Two clones would share a cursor and range; prevent that activation path or add an external/non-clonable allocation boundary.
What the answer must demonstrate: Discard its entire old remainder and request a fresh disjoint range under a new incarnation.
A replica missing a committed high-water advance can hand out values already reserved elsewhere. Preserve acknowledged range state and request results through the chosen promotion protocol. A stale backup needs equivalent reconciliation or refusal; restarting from its counter silently is unsafe.
Interviewer follow-up
Does an order unique constraint solve this?
Reveal the follow-up answer
It detects collisions but does not make a violated allocation contract correct.
What the answer must demonstrate: A replica missing a committed high-water advance can hand out values already reserved elsewhere.
Some clients cannot exactly represent all supported integers. JavaScript Number loses exactness above 2^53 − 1, so parsing through it can merge distinct IDs. A decimal string and appropriate integer type preserve the value and its namespace.
Interviewer follow-up
Can numeric sort prove creation order?
Reveal the follow-up answer
No. Different range holders progress independently; use explicit business time or a separate ordering service.
What the answer must demonstrate: Some clients cannot exactly represent all supported integers.
When compact approximate time ordering is an actual requirement. Timestamp, worker and sequence fields introduce clock rollback, worker reuse and overflow rules that numeric ranges avoid. UUID standards are another alternative when 128-bit storage is acceptable; strict global order still needs coordination.
Interviewer follow-up
Can a sequence provide gapless invoices?
Reveal the follow-up answer
Not by itself. Numbering must follow the business commit/cancellation policy.
What the answer must demonstrate: When compact approximate time ordering is an actual requirement.
Under this contract it means the event and processing work were durably accepted. The receiver can respond before completing its business workflow, which preserves low endpoint latency without losing work on process restart. A sender needing proof of completed processing requires a separate status or callback protocol.
Interviewer follow-up
Why not acknowledge an in-memory queue?
Reveal the follow-up answer
A crash could erase the event after the sender stopped retrying.
What the answer must demonstrate: Under this contract it means the event and processing work were durably accepted.
Event E402, delivery D22 and immutable body remain stable. Attempt identity, timestamp, signature and sender lease token change. This lets diagnostics distinguish network attempts while the receiver recognizes the same business event. A new event ID would undermine deduplication.
Interviewer follow-up
Why include endpoint ID in planning uniqueness?
Reveal the follow-up answer
Configuration version numbers are local to endpoints; two endpoints at version 3 both need their own delivery.
What the answer must demonstrate: Event E402, delivery D22 and immutable body remain stable.
It cannot infer acceptance from the timeout. Refusing every retry loses events in the history where the request never arrived; retrying can repeat arrival when receipt committed. Receiver inbox identity makes that repetition safe under the agreed contract, or a supported receipt lookup can resolve it.
A1 reports timeout after A2 succeeds. What changes?
Reveal a model answer
The delivery row changes only if the result update carries the current lease token and expected in-flight state. A stale A1 cannot overwrite A2’s accepted state. Its observation can be retained separately in attempt history. This local fence does not prevent remote duplicate POSTs.
Interviewer follow-up
What handles those duplicates?
Reveal the follow-up answer
The receiver’s trusted inbox and business-effect protocol.
What the answer must demonstrate: The delivery row changes only if the result update carries the current lease token and expected in-flight state.
Version 8 arrives before version 7. May version 7 be ignored?
Reveal a model answer
If events carry complete snapshots, a version guard can safely retain the newer state. If events are dependent deltas, dropping version 7 may lose a necessary effect; detect gaps and replay or fetch authoritative complete state. Waiting for each 202 orders acceptance, not necessarily receiver processing.
Interviewer follow-up
Can timestamps replace an ordering token?
Reveal the follow-up answer
Not generally. Clock and delivery order can differ, and timestamps can tie.
What the answer must demonstrate: If events carry complete snapshots, a version guard can safely retain the newer state.
How does a five-second failing endpoint affect capacity?
Reveal a model answer
At the retry-adjusted peak, thousands of attempts per second can occupy more than 20,000 sockets. Per-endpoint concurrency and tenant scheduling shares prevent one receiver using the whole pool; global limits protect the fleet. Backoff retains obligations without repeatedly hammering the destination.
Interviewer follow-up
Why not scale workers without those limits?
Reveal the follow-up answer
That amplifies load and cost while violating receiver and fairness budgets.
What the answer must demonstrate: At the retry-adjusted peak, thousands of attempts per second can occupy more than 20,000 sockets.
A URL can resolve to a different destination after registration. Validate the actual chosen address and connect to it while checking TLS for the original hostname. Otherwise DNS rebinding or an unchecked second resolution can reach internal resources. Redirects require the same validation or must be disabled.
Interviewer follow-up
Does signature verification make the URL safe?
Reveal the follow-up answer
No. Payload authenticity and outbound destination restrictions protect different boundaries.
What the answer must demonstrate: A URL can resolve to a different destination after registration.
Can an operator safely replay a three-month-old event?
Reveal a model answer
This design retains payloads for seven days, so it rejects unavailable history. A longer archive contract must preserve original bytes and destination ownership, and receiver deduplication or business reconciliation must cover that horizon. Reconstructing the current object is a new event, not faithful replay.
Interviewer follow-up
What if an external effect key expired?
Reveal the follow-up answer
Do not blindly resend; use supported lookup or explicit reconciliation before creating another effect.
What the answer must demonstrate: This design retains payloads for seven days, so it rejects unavailable history.
At the assumed billion evaluations/s, network calls become enormous work and add a service dependency to application latency. Local compiled snapshots make lookups cheap and let requests continue briefly during control outages. That benefit requires an explicit stale-use and fallback policy rather than claiming instant global updates.
Interviewer follow-up
When would you choose remote evaluation?
Reveal the follow-up answer
When current authority or confidential rule access justifies the network and availability cost.
What the answer must demonstrate: At the assumed billion evaluations/s, network calls become enormous work and add a service dependency to application latency.
Why does expanding ten percent to twenty preserve tenant 54?
Reveal a model answer
Its bucket 731 remains below both thresholds 1,000 and 2,000 when seed, targeting key, encoding and hash algorithm stay unchanged. A new random draw each request would not preserve assignment. Tenant-based targeting gives all authorized users of that tenant the same cohort.
Interviewer follow-up
What does a seed change do?
Reveal the follow-up answer
It deliberately changes hash inputs and can reshuffle the cohort; treat it as a migration.
What the answer must demonstrate: Its bucket 731 remains below both thresholds 1,000 and 2,000 when seed, targeting key, encoding and hash algorithm stay unchanged.
A request can still read F7 after its update and F8 before its dependent update, creating a configuration never validated together. Compile a complete immutable snapshot and retain one reference throughout related evaluations. Atomic pointer replacement changes later requests while old references safely finish.
Interviewer follow-up
When may old compiled data be freed?
Reveal the follow-up answer
Only after no active request or retained policy reference still needs it.
What the answer must demonstrate: A request can still read F7 after its update and F8 before its dependent update, creating a configuration never validated together.
C18 installs before a slow C17 download. What happens?
Reveal a model answer
The installer validates bytes then compares generations under an atomic update. Since 17 is older than 18, it is ignored. A real rollback also uses a higher generation containing the desired older values. Thus content reversal never requires reversing publication order.
Interviewer follow-up
What if two installers read the same old pointer?
Reveal the follow-up answer
Compare-and-swap lets one succeed; the other rechecks against the changed pointer.
What the answer must demonstrate: The installer validates bytes then compares generations under an atomic update.
Why cannot a successful C17 cache fetch renew freshness?
Reveal a model answer
A cache can return authentic C17 after C18 was published. Periodic synchronization must check the environment publication authority, not merely download old bytes again. The main design tracks synchronization and falls back after a long outage; it does not claim a strict publication-to-disable deadline.
Interviewer follow-up
Does successful synchronization prove every instance installed the current generation?
Reveal the follow-up answer
No. It describes that synchronization, not the entire fleet. Observe active-generation distribution and actual rollout adoption across instances.
What the answer must demonstrate: A cache can return authentic C17 after C18 was published.
What happens when an instance cannot hear the off switch?
Reveal a model answer
The application uses its last validated snapshot during a short outage, then returns the documented false fallback after five minutes without successful synchronization. That is an outage policy, not an instantaneous command or a proved global five-minute disable bound. Sensitive actions need their own current permission check.
Interviewer follow-up
What happens after restart with saved C17?
Reveal the follow-up answer
Require a new authority confirmation; do not reset old bytes to a fresh age.
What the answer must demonstrate: The application uses its last validated snapshot during a short outage, then returns the documented false fallback after five minutes without successful synchronization.
No. The decision must influence the experience under the agreed exposure definition. Debug calls, hidden branches and requests that never render can evaluate without exposing a variant. Record actual flag, generation and variant used, and analyze outcome guardrails in addition to adoption.
Interviewer follow-up
Does a percentage allocation prove causal improvement?
Reveal the follow-up answer
No. It is one mechanism within an experiment requiring valid assignment, observation and analysis.
What the answer must demonstrate: No. The decision must influence the experience under the agreed exposure definition.
Run identical fixed context/seed inputs against expected bucket and rule outputs in every language. Test type and missing-attribute behavior, snapshot installation order, offline expiry and restart. Observe generation spread and fallback rates during a limited rollout. Rule schema/compiler compatibility must be checked before activation.
Interviewer follow-up
Can a flag undo an incompatible database migration?
Reveal the follow-up answer
No. Data and code compatibility need their own staged migration and recovery plan.
What the answer must demonstrate: Run identical fixed context/seed inputs against expected bucket and rule outputs in every language.
It expresses the requested organization, not membership or permission. Authentication establishes the actor; authorization verifies the active tenant and action. Trusted scope then travels through database queries, caches, jobs and files. An unchecked forwarded tenant header would let a caller choose another customer’s data.
Interviewer follow-up
What key identifies an invoice?
Reveal the follow-up answer
Tenant plus local invoice ID. Related records and cache keys must include the same tenant.
What the answer must demonstrate: It expresses the requested organization, not membership or permission.
It lets the database enforce row-access policy as defense in depth. Its assumptions include a constrained application role and initialized transaction-local tenant context. Table owners, bypass privileges or stale pooled-session context can undermine those assumptions. It complements rather than replaces application authorization.
Why does request-count limiting miss export overload?
Reveal a model answer
A million-row export can consume roughly 100 CPU-seconds in the example, equivalent to 50,000 ordinary reads. Counting it as one request hides database work and output load. Limit active exports, query duration, scanned/output bytes and connections, with fair tenant scheduling and reserved interactive capacity.
Use measured sustained demand or explicit region, key, recovery and administration constraints. Ordinary tenants can pool economically in cells. Separate schemas mainly organize namespaces; dedicated compute is needed to isolate certain resource contention. Every placement still requires logical authorization at its access paths.
Interviewer follow-up
What remains shared across cells?
Reveal the follow-up answer
Identity, directory and control services can remain common dependencies requiring separate availability and rollout controls.
What the answer must demonstrate: Use measured sustained demand or explicit region, key, recovery and administration constraints.
The chosen migration pauses tenant writes at a boundary shared by every mutation. It waits for admitted writes to commit or abort before copying the stable data. A writer that completed first is copied with its replay result; one arriving after the freeze is rejected and retries at the new placement. Background writers must follow the same rule.
Interviewer follow-up
Which overlooked path can break this?
Reveal the follow-up answer
Any background/admin mutation that bypasses the same guard, including job completion after a move.
What the answer must demonstrate: The chosen migration pauses tenant writes at a boundary shared by every mutation.
No. A timeout does not prove C5 accepted no writes. Keep C2 frozen, inspect durable migration progress and recover or finish the destination transition. Once C5 has new writes, returning to C2 needs a controlled transfer of those changes. The main design accepts a maintenance window rather than simultaneous ambiguous ownership.
Interviewer follow-up
Why not just change the directory back?
Reveal the follow-up answer
After destination writes, the old source is stale. Safe return requires a new transfer or destination repair.
What the answer must demonstrate: No. A timeout does not prove C5 accepted no writes.
Does a signed download URL enforce current user identity?
Reveal a model answer
The selected design uses an authenticated gateway that checks current tenant membership and permission on a new download request. An ordinary presigned URL is an alternative bearer capability: possession may allow access until expiry even after revocation. Use that alternative only if its expiry-bound behavior meets the contract.
Interviewer follow-up
What must the worker publish?
Reveal the follow-up answer
The exact immutable object version it validated, with tenant ownership and current completion authority.
What the answer must demonstrate: The selected design uses an authenticated gateway that checks current tenant membership and permission on a new download request.
How do you restore one tenant from a pooled backup?
Reveal a model answer
Restore into a separate recovery environment, verify tenant scope and import only intended records through a controlled ownership plan. Include replay outcomes, outbox and placement-control state, then rebuild derived views. Restoring the pooled database in place would overwrite other customers.
Interviewer follow-up
What test demonstrates isolation better than a query unit test?
Reveal the follow-up answer
Run two tenants with matching local IDs through cache, pooled connections, jobs and file APIs, then attempt cross-scope access.
What the answer must demonstrate: Restore into a separate recovery environment, verify tenant scope and import only intended records through a controlled ownership plan.
The key is the mutable name users request; a version identifies exact immutable bytes. Replacement changes the pointer while an existing reader finishes on its selected version. This prevents mixed downloads.
Interviewer follow-up
Does a key prefix authorize access?
Reveal the follow-up answer
No. Verify actor, tenant and operation separately at metadata and delivery boundaries.
What the answer must demonstrate: The key is the mutable name users request; a version identifies exact immutable bytes.
What if the process crashes after bytes are written but before publication?
Reveal a model answer
The bytes remain unreferenced temporary work, and the logical object has not changed. A retry can finish the session or cleanup can later reclaim it. Reversing the order could publish a pointer to missing bytes.
Interviewer follow-up
What if the reply is lost after publication?
Reveal the follow-up answer
The completed session records the version and returns that same result on retry.
What the answer must demonstrate: The bytes remain unreferenced temporary work, and the logical object has not changed.
It creates 32 independently transferable pieces. The client retries a failed part and can use bounded parallelism. Completion verifies the ordered immutable parts and publishes their manifest without copying the entire video in a transaction.
Interviewer follow-up
Can part replacement mutate a published version?
Reveal the follow-up answer
No. Store a new immutable part identity and preserve the identity referenced by the published manifest.
What the answer must demonstrate: It creates 32 independently transferable pieces. The client retries a failed part and can use bounded parallelism.
The metadata authority conditionally updates the key. The first committed replacement wins; the other receives a conflict when V4 is no longer current. This is an explicit lost-update policy rather than accidental last-response-wins behavior.
Interviewer follow-up
Can unconditional PUT still be supported?
Reveal the follow-up answer
Yes, if its last-committed-writer semantics are explicit and the caller chooses them.
What the answer must demonstrate: The metadata authority conditionally updates the key.
After all three required independent byte copies are durable and the synchronously replicated metadata publication commits. If a required copy cannot be written, wait or fail completion without publishing it. Merely queueing replication cannot promise survival immediately after success.
Only if placement actually separates the relevant disks, hosts or zones.
What the answer must demonstrate: After all three required independent byte copies are durable and the synchronously replicated metadata publication commits.
V5 is published midway through a V4 range download. What happens?
Reveal a model answer
The existing download continues against the exact V4 manifest and immutable bytes. A new current-key lookup can select V5. Cleanup retains versions needed by active or retained readers.
Does revoking membership immediately invalidate a signed URL?
Reveal a model answer
Not necessarily. An ordinary signed URL is a bearer grant until its effective expiry. Short expiry limits exposure; an authenticated gateway can enforce current permission on new delivery requests. Neither recalls bytes already received.
A completed version may reference that part, and a retained or active reader may still require it. Cleanup must distinguish expired unreferenced uploads from published data. Conservative retention is the simple starting point.
Why can the first implementation use ordinary text search?
Reveal a model answer
A relational text index can search five hundred policies. The application retrieves passages, checks current permissions and generates cited answers without a separate vector service. Add semantic search when tests show keyword search misses relevant passages because questions use different wording.
How can a chunking decision make the taxi answer wrong?
Reveal a model answer
If one chunk contains the reimbursement rule and another contains its manager-approval condition, retrieving only the first gives the generator incomplete evidence. Preserve related conditions where possible and include evaluation questions that require them.
Interviewer follow-up
Why not use the entire policy document?
Reveal the follow-up answer
It increases context cost and can bury the relevant rule among exceptions or unrelated sections. Choose chunks from measured retrieval and answer quality, with a bounded total context.
What the answer must demonstrate: Connect chunk boundaries to the missing approval condition, rather than treating chunk size as an arbitrary tuning constant.
Use tenant filters during search, then check current grants, source version and deletion state before each model receives private passages, including a reranker. Recheck sources before release and when opening citations. Authorize private answer-object access separately; source access does not grant access to another employee’s question.
Interviewer follow-up
Does that guarantee a revocation stops all in-flight text immediately?
Reveal the follow-up answer
No. The catalog decision orders admission against grant changes. A revocation after admission cannot undo processing already admitted or bytes already delivered; stronger guarantees need an explicit additional protocol and still cannot erase prior recipients.
What the answer must demonstrate: Check private answer-object access separately from evidence access; authorize each private-text recipient and state the in-flight limit.
What happens when version 13 is current but only version 12 is indexed?
Reveal a model answer
The selected contract rejects version 12 for new answers. The assistant may temporarily have insufficient current evidence while indexing catches up. It must not silently represent the older policy as current.
Interviewer follow-up
Could a product deliberately allow older evidence?
Reveal the follow-up answer
Yes, with an explicit freshness policy and visible version information, while retaining current authorization. That is a different product contract, not an invisible implementation shortcut.
What the answer must demonstrate: Separate the current-version contract from index readiness and describe the resulting temporary evidence gap.
It proves only that the reference identifies a supplied source. The answer may still reverse its meaning, omit a condition or combine incompatible statements. Validate citation identity and evaluate claim support and task correctness separately.
Interviewer follow-up
What should the taxi test assert?
Reveal the follow-up answer
That P7 version 12 is the cited evidence and the answer preserves the manager-approval condition, not merely that some citation is present.
What the answer must demonstrate: Distinguish reference identity from claim support; preserve the manager-approval qualification in the example.
The source-version transaction records durable indexing work. A worker retries it with stable document/version/chunk identities, so it can complete missing writes without duplicating passages. The query path still rejects deleted or superseded sources.
Interviewer follow-up
Why is a successful document update not proof of search readiness?
Reveal the follow-up answer
The authoritative version can commit before asynchronous extraction and indexing finish. Search readiness and catalog durability are separate states.
What the answer must demonstrate: Use durable follow-up work and stable versioned chunk identities; do not equate source commit with search readiness.
Which capacity estimate matters after search becomes fast?
Reveal a model answer
The model token workload. At 200 requests per second and 4,800 input tokens each, peak input demand is 960,000 tokens per second; 400 output tokens add 80,000 per second. Bound context, output and concurrency rather than relying on request count alone.
Interviewer follow-up
How should a bulk reindex share model capacity?
Reveal the follow-up answer
Give background embedding work a separate budget so it cannot consume the interactive capacity promised to employee questions.
What the answer must demonstrate: Calculate input and output token demand and protect interactive capacity from background ingestion.
What should users receive during search, permission and model failures?
Reveal a model answer
If permissions cannot be checked, deny private access. If search is down, report retrieval unavailability rather than claim no evidence exists. If generation is down, currently authorized excerpts can be returned as an explicitly labeled fallback.
Interviewer follow-up
What is the most useful next quality experiment?
Reveal the follow-up answer
Compare keyword-only and hybrid retrieval on labeled questions, measuring candidate recall and supported-answer correctness separately from latency and cost.
What the answer must demonstrate: Distinguish unavailable authority or retrieval from absent evidence, and keep excerpt fallbacks currently authorized.
What are prefill and decode, and why measure them separately?
Reveal a model answer
Prefill processes input tokens and builds cached attention state. Decode repeatedly generates the next token using the prefix and cache. A long prefill can delay existing decoders, so first-token latency and gaps between output tokens reveal different scheduling problems.
Do isolated bounds of 27 prefill workers and 40 decode workers prove forty workers are enough?
Reveal a model answer
No. The isolated measurements each assume a particular workload, while mixed workers share compute, memory bandwidth and capacity between phases. Use those figures as lower bounds, then load-test the combined prompt/output distribution with headroom and latency targets.
Interviewer follow-up
Why calculate active streams?
Reveal the follow-up answer
Each live sequence occupies KV state. At 100 arrivals per second and about twelve seconds of decode, roughly 1,200 sequences are active before queue and prefill time are included.
What the answer must demonstrate: Treat isolated throughput as lower bounds, then account for mixed execution and live-sequence memory.
Two gateways see the same worker with one free slot. What prevents over-admission?
Reveal a model answer
The worker makes a guarded reservation against its actual sequence and KV capacity before accepting execution. Registry reports guide routing but do not allocate memory. One reservation succeeds; the other must wait within its deadline or use another eligible worker before execution starts.
Interviewer follow-up
Why start with maximum-length reservation?
Reveal the follow-up answer
It provides a straightforward memory bound using the input plus allowed output length. Dynamic allocation may use memory better but needs explicit handling when sequences grow and capacity runs out.
What the answer must demonstrate: Put the guarded allocation at the worker; explain the simplicity and utilization cost of maximum-length reservation.
How do continuous batching and chunked prefill solve different problems?
Reveal a model answer
Continuous batching lets finished sequences leave and new ones enter between iterations, avoiding empty fixed-batch slots. Chunked prefill limits how much long-prompt work runs at once so active decoders get opportunities to produce output.
Yes. Larger batches can lengthen token gaps, and strong decode priority can delay new prefills. Measure both first-token and inter-token latency for the actual workload.
What the answer must demonstrate: Distinguish dynamic batch membership from prefill scheduling and measure both latency consequences.
Why is it unsafe to free the KV cache as soon as a client disconnects?
Reveal a model answer
A GPU kernel may still be reading those blocks. Immediate reuse by another request can corrupt live execution. Record cancellation, stop future scheduling, wait for in-flight references to finish, and only then release blocks.
Interviewer follow-up
What if the worker is unreachable?
Reveal the follow-up answer
Record cancellation intent and rely on the bounded generation deadline while establishing worker loss. Do not claim physical work stopped instantly; status may become interrupted rather than acknowledged canceled.
What the answer must demonstrate: Preserve in-flight memory references and distinguish recorded cancellation intent from confirmed stopped execution.
Exact effective prefix tokens, compatible model weights/tokenizer/template/adapters, and a server-controlled tenant or approved trust scope. Shared blocks remain immutable while live requests reference them. A cache entry is computation state, not an independently authorized answer.
Interviewer follow-up
Why not let clients choose the tenancy salt?
Reveal the follow-up answer
A client could select another tenant's scope and create cross-tenant reuse or timing exposure. The gateway derives isolation settings from authenticated identity.
What the answer must demonstrate: Include exact compatible prefix identity and trusted tenant scope, with immutable shared live state.
What survives a worker crash in the selected design?
Reveal a model answer
The generation identity, status and recorded usage survive in durable storage. Live KV tensors and unsaved output events disappear. Mark the dead process’s generations interrupted. A restarted worker has a new incarnation—an identity for that process start—and rejects old assignments; a replacement generation is an explicit new attempt.
Interviewer follow-up
Why is saving only the prompt insufficient?
Reveal the follow-up answer
It can start another computation, but does not restore the exact prior execution state or guarantee the same sampled continuation.
What the answer must demonstrate: Separate durable identity from ephemeral state, bind assignments to process incarnations, and make replacement attempts explicit.
How do cumulative usage reports avoid duplicate charges?
Reveal a model answer
Advance the saved count monotonically for one generation: reports of 100 and 120 output tokens yield 120, not 220. Finalize a terminal count so late reports cannot reopen billing. Under this design, work after the last durable report at a crash is an internal unbilled cost.
Output quality, tokenizer/template compatibility, mixed-load first-token and inter-token latency, memory, cancellation and recovery. New workers load the tested version while old workers drain.
What the answer must demonstrate: Advance cumulative counts rather than adding them, finalize terminal totals, and disclose unrecorded crash-tail work.
What makes this more reliable than replaying a chat transcript?
Reveal a model answer
The service stores explicit workflow state, bounded activity results, exact proposals, authenticated approvals and external action identities. A restarted worker resumes from those facts instead of asking the model to reconstruct what happened.
Interviewer follow-up
Can the model still choose tools?
Reveal the follow-up answer
It can propose permitted operations inside a bounded workflow, but trusted code validates and authorizes execution.
What the answer must demonstrate: The service stores explicit workflow state, bounded activity results, exact proposals, authenticated approvals and external action identities.
Do 2,083 tasks waiting for approval require 2,083 workers?
Reveal a model answer
No. Waiting tasks are durable records with timers. Workers run ready activities and release capacity while approval is pending. Worker sizing follows active call demand and latency rather than the number of open tasks.
Interviewer follow-up
What wakes a waiting task?
Reveal the follow-up answer
A validated approval, cancellation or expiry transition makes the appropriate next step ready.
What the answer must demonstrate: No. Waiting tasks are durable records with timers.
P8 was approved, but the vendor changes the price. May the agent submit?
Reveal a model answer
Not under the old approval. Validate the saved quote and exact proposal before admission; changed commercial terms require a new proposal and approval. Generated text saying the new price is acceptable does not authorize it.
Interviewer follow-up
What fields should approval bind?
Reveal the follow-up answer
Vendor, items, quantity, amount, currency, destination, validity and the immutable proposal identity.
What the answer must demonstrate: Not under the old approval. Validate the saved quote and exact proposal before admission; changed commercial terms require a new proposal and approval.
The worker crashes after saving a model result. Should it call the model again?
Reveal a model answer
It should reuse the saved activity result and resume the next state. If the result was never committed, repeating a bounded read or generation may be acceptable; external purchases require a different recovery contract.
Interviewer follow-up
Why preserve workflow versions?
Reveal the follow-up answer
A new deployment should not silently reinterpret saved state or reorder actions in an existing task.
What the answer must demonstrate: It should reuse the saved activity result and resume the next state.
The vendor accepted an order but its response was lost. What next?
Reveal a model answer
Recover the same action identity through the provider’s documented idempotent retry or lookup within its supported retention window. Save the confirmed order if found. If it cannot be resolved, report SubmissionUnknown and reconcile; a new key risks another order.
No. It reliably dispatches intended work but may dispatch again; the external effect still needs identity and recovery.
What the answer must demonstrate: Recover the same action identity through the provider’s documented idempotent retry or lookup within its supported retention window.
Why is a worker lease insufficient to guarantee one purchase?
Reveal a model answer
An expired worker may already have sent the request, and a remote provider can act after local ownership changes. Leases and task versions coordinate local progress; a stable external action key handles duplicate submission where the provider supports it.
Interviewer follow-up
What if the provider supports neither idempotency nor lookup?
Reveal the follow-up answer
Do not claim safe automatic recovery of an uncertain action; keep it unresolved for controlled reconciliation.
What the answer must demonstrate: An expired worker may already have sent the request, and a remote provider can act after local ownership changes.
Can cancel always promise that no order was created?
Reveal a model answer
Only before submission has been admitted and sent. Afterward the action may already exist remotely. Resolve that action and use the vendor’s cancellation or compensation process under a tracked identity. The UI must distinguish these states.
Interviewer follow-up
Can changing the proposal cancel the old one implicitly?
Reveal the follow-up answer
No. Track any admitted action separately and resolve it; new proposal state does not erase an external request.
What the answer must demonstrate: Only before submission has been admitted and sent.
A retrieved vendor document tells the model to send credentials elsewhere. What stops it?
Reveal a model answer
Retrieved text is untrusted data. Tool gateways allow only validated operations, enforce current permissions and keep secrets in trusted integrations. The model cannot authorize a destination or grant itself a credential.
Interviewer follow-up
What operational signal most deserves attention?
Reveal the follow-up answer
Unresolved external submissions and approval mismatches deserve direct inspection even when overall completion rates look healthy.
What the answer must demonstrate: Retrieved text is untrusted data. Tool gateways allow only validated operations, enforce current permissions and keep secrets in trusted integrations.
Why is regional popularity a valid starting design?
Reveal a model answer
It provides useful context-sensitive results without historical personalization or a learned model. The complete flow still filters eligibility, limits repeated creators, records response identity and collects actual feedback. It becomes both a comparison baseline and an outage fallback.
Interviewer follow-up
What limitation motivates personalization?
Reveal the follow-up answer
People in the same region can have different interests, so popularity may miss relevant niche content.
What the answer must demonstrate: It provides useful context-sensitive results without historical personalization or a learned model.
Why not score all ten million videos on every request?
Reveal a model answer
At five thousand requests/s that is fifty billion scores/s. At the illustrative 50 microseconds each, it needs about 2.5 million busy cores before other work. Retrieval narrows the pool so richer ranking is affordable.
Interviewer follow-up
What does two hundred candidates change?
Reveal the follow-up answer
It reduces peak scoring to one million scores/s, roughly fifty busy cores at that assumed cost, before headroom and feature work.
What the answer must demonstrate: At five thousand requests/s that is fifty billion scores/s.
What is the difference between candidate retrieval and ranking?
Reveal a model answer
Retrieval finds a limited set likely to contain useful items using sources such as followed creators, topics, similarity and popularity. Ranking spends more information and computation comparing that set. A perfect ranker cannot select a relevant item retrieval never supplied.
Interviewer follow-up
How do new items enter?
Reveal the follow-up answer
Use metadata-based candidates or bounded exploration rather than requiring watch history they do not yet have.
What the answer must demonstrate: Retrieval finds a limited set likely to contain useful items using sources such as followed creators, topics, similarity and popularity.
I12 has a high score but U7 completed it. What happens?
Reveal a model answer
Under this product contract it is filtered out. Eligibility and completion rules are explicit constraints, not tiny penalties a sufficiently high engagement score can overcome. Diversity rules then shape the eligible ranked page.
Interviewer follow-up
Why can the final list differ from pure score order?
Reveal the follow-up answer
Creator caps and diversity can improve the overall experience while lowering the sum of individual predicted scores.
What the answer must demonstrate: Under this product contract it is filtered out.
Why is returning an item not enough to label it as ignored?
Reveal a model answer
The viewer may never have seen it. Record actual visibility separately from the returned list and use labels appropriate to the observation. Unknown exposure is not the same as a negative preference.
Interviewer follow-up
How do duplicated events avoid inflating popularity?
Reveal the follow-up answer
Reuse stable event IDs and couple deduplication with aggregate updates for the declared replay horizon.
What the answer must demonstrate: The viewer may never have seen it.
Training or evaluation uses facts that were not available at the historical decision, such as later popularity or a future watch. Offline metrics then overstate what serving could have achieved. Use historically appropriate feature values and availability.
Interviewer follow-up
Does good offline ranking prove product improvement?
Reveal the follow-up answer
No. Use a controlled online rollout with satisfaction, safety and latency guardrails.
What the answer must demonstrate: Training or evaluation uses facts that were not available at the historical decision, such as later popularity or a future watch.
The personalized feature store times out. What should the API return?
Reveal a model answer
Use a tested simpler score or contextual popularity within the latency budget, while retaining eligibility checks. Expose feature age and fallback rate to operations. A relevance outage should not become a reason to serve forbidden content.
Interviewer follow-up
What if eligibility cannot be checked?
Reveal the follow-up answer
Omit uncertain candidates or use a pool whose eligibility can be established under the declared policy.
What the answer must demonstrate: Use a tested simpler score or contextual popularity within the latency budget, while retaining eligibility checks.
Stop using personal history for candidate retrieval and ranking, invalidate personalized continuation state and serve contextual results. Handle stored and derived history under the documented deletion policy and propagation limits.
Interviewer follow-up
Why not put raw user IDs into every metric label?
Reveal the follow-up answer
It creates sensitive, high-cardinality monitoring data; retain controlled investigation paths and aggregate operational metrics.
What the answer must demonstrate: Stop using personal history for candidate retrieval and ranking, invalidate personalized continuation state and serve contextual results.
Why can signaling succeed while the call has no media?
Reveal a model answer
Signaling only exchanges control information. Media still needs compatible settings, a reachable direct or relayed network path and successful encrypted transport. Firewalls or relay exhaustion can block media even when the application socket works.
Interviewer follow-up
What should join success measure?
Reveal the follow-up answer
Time to actual usable received audio or video, alongside the control-plane milestones.
What the answer must demonstrate: Signaling only exchanges control information. Media still needs compatible settings, a reachable direct or relayed network path and successful encrypted transport.
ICE gathers and tests candidate paths. STUN helps discover the public address observed outside a local network. TURN relays media when a usable direct path is unavailable. Discovery alone does not supply relay capacity.
Interviewer follow-up
Why test TURN specifically?
Reveal the follow-up answer
A design tested only on friendly direct networks may fail behind restrictive networks or when relay capacity is exhausted.
What the answer must demonstrate: ICE gathers and tests candidate paths. STUN helps discover the public address observed outside a local network.
What changes when six 1.5 Mbps publishers move from mesh to an SFU?
Reveal a model answer
In mesh each uploads five copies, or 7.5 Mbps. With one stream to the SFU each uploads 1.5 Mbps and SFU ingress is 9 Mbps. Full-quality forwarding to everyone still produces 45 Mbps of server egress.
Interviewer follow-up
How does subscription selection help?
Reveal the follow-up answer
One 1.5 Mbps speaker plus four 0.15 Mbps tiles uses 2.1 Mbps per receiver, reducing the room’s selected output.
What the answer must demonstrate: In mesh each uploads five copies, or 7.5 Mbps.
Why would a publisher deliberately upload several qualities?
Reveal a model answer
Simulcast gives the SFU ready-made quality choices for different receivers. It raises publisher encoding work and upload but lets a weak receiver or small tile receive less data without server-side transcoding for every subscription.
Interviewer follow-up
Does more buffering solve congestion?
Reveal the follow-up answer
No. It adds conversational delay; adapt bitrate and drop obsolete video while protecting audio.
What the answer must demonstrate: Simulcast gives the SFU ready-made quality choices for different receivers.
The room service saves the authorized decision, and the SFU stops P6’s publication and subscriptions. Reconnect checks current membership. A browser roster update alone does not remove access to forwarded media.
Interviewer follow-up
Can notification delivery prove instantaneous removal?
Reveal the follow-up answer
No. Target five seconds on healthy paths, reconcile every fifteen seconds, and stop affected sessions after sixty seconds without successful authority refresh. These are a target and outage policy; a proved global worst-case deadline needs stronger timing and authority assumptions.
What the answer must demonstrate: The room service saves the authorized decision, and the SFU stops P6’s publication and subscriptions.
Clients establish fresh transports to a replacement SFU that loads current room permission and subscriptions. There is an interruption; dead in-flight packets and transport state are not transparently restored.
Interviewer follow-up
What if only signaling disconnects?
Reveal the follow-up answer
Existing media can continue under the current session policy, but joins and membership changes are impaired and permission freshness still matters.
What the answer must demonstrate: Clients establish fresh transports to a replacement SFU that loads current room permission and subscriptions.
Is a call automatically end-to-end encrypted because transport is encrypted?
Reveal a model answer
No. Encrypted browser-to-SFU and SFU-to-browser legs may still let the trusted SFU access media. Encryption that excludes the server changes recording, transcription and moderation capabilities and needs its own key-management design.
Interviewer follow-up
How should recording be introduced?
Reveal the follow-up answer
As an explicitly authorized media subscriber with consent, retention, access and encryption behavior defined.
What the answer must demonstrate: No. Encrypted browser-to-SFU and SFU-to-browser legs may still let the trusted SFU access media.
Which load test is more useful than opening many signaling sockets?
Reveal a model answer
Realistic packet traffic with subscriptions, loss, bitrate adaptation, relay paths and region capacity reveals the media bottlenecks. Measure audio gaps, video freezes, join delay and reconnect time as well as CPU and egress.
Interviewer follow-up
What resource can saturate with moderate CPU?
Reveal the follow-up answer
Network egress or packet-processing capacity can limit an SFU before raw compute utilization looks high.
What the answer must demonstrate: Realistic packet traffic with subscriptions, loss, bitrate adaptation, relay paths and region capacity reveals the media bottlenecks.