System-design interview · Extended interviews
Design an LLM inference platform
Serve versioned language models with bounded queues, token-aware admission and safe streaming; use measured scheduling and memory improvements before introducing specialized distributed execution.
You will learn to
- Explain prefill, decode and KV memory through one complete generation.
- Derive token throughput and memory budgets, then choose scheduling and scaling changes.
- Handle cancellation, worker loss and usage reporting without pretending lost execution state can be resumed.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Capacity estimation: throughput, latency, concurrency and storage · Message queues, event logs, delivery guarantees, and backpressure · Authentication, authorization, and tenant isolation
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Define the generation service and its limits
Build a service that accepts an authenticated text prompt, generates output from a selected model version and streams that output to the caller. Include tenant quotas, cancellation and model rollouts. Exclude model training and execution of generated tool calls. A proposed tool call is output data; another application decides whether it may run.
Ask whether output must stream, which prompt/output sizes matter, and whether a stream must continue seamlessly after worker failure. Here output streams, the worked request has 4,000 input tokens and at most 600 output tokens, and worker loss may interrupt it.
Follow generation G81: a 4,000-token prompt and an upper limit of 600 output tokens. Tokens are units from the model's tokenizer, often word fragments. The service must budget this request differently from a twenty-token question, even though both count as one HTTP request.
Choose an explicit recovery contract. Request identity and status are durable; live model execution state is not. Worker failure interrupts active generations. A client can start a new generation, but the service does not silently append a different regenerated answer to the old partial stream.
02Functional requirements
Agree on what the service must do before choosing its components.
- Generate versioned output. Accept an authenticated prompt and selected immutable model version, then stream numbered output events with a visible terminal status.
- Inspect and retry a request. Return a generation identity and status. Repeating the same tenant/request identity and effective settings retrieves that logical generation; conflicting settings fail.
- Cancel generation. Accept an idempotent cancellation request and show whether execution has stopped, completed or been interrupted. Recording a cancel request does not prove the worker has already stopped.
- Report usage and roll out models. Expose recorded token usage, apply tenant quotas and deploy tested model versions while draining old requests. Training and execution of generated tool calls are outside this service.
03Non-functional requirements
Use these hypothetical requirements for the worked interview. Confirm the assumptions with the interviewer; the numerical targets require testing and are not measured results. For latency, p95 and p99 mean that 95% and 99% of measured delays, respectively, are no greater than the reported value.
- Workload. Size the illustrative load at 100 requests/s with 4,000 input and 600 output tokens per request: 400,000 input and 60,000 output tokens/s. Validate the actual mix of shorter/longer requests separately rather than assuming every request costs the same.
- Interactive latency. Target p95 time to first token below one second and p95 gaps between subsequent tokens below 50 ms for the tested input/output distribution. Queue waiting consumes the first-token budget; throughput alone does not validate either target.
- Durable control, interruptible execution. Request identity, status and recorded usage survive the supported single control-database-node failure through durable majority commits. Live worker state is temporary: worker loss interrupts active generations, and lost database majority stops new admissions. Do not claim transparent continuation or instantaneous cancellation through a partition.
- Bounded, fair work. Limit tenant token work, concurrency, queue time, output length and absolute deadlines. Admit only work fitting the worker’s actual memory budget; slow clients receive bounded buffers and cannot consume capacity forever.
- Private, compatible execution. Authorize generation status/stream ownership and scope private input reuse to trusted tenants. Pin compatible weights, tokenizer, prompt template and runtime; protect artifacts and define private-input retention. Generated output is data, not permission to execute server code.
- Consistent accounting. Persist cumulative usage monotonically and finalize a terminal total. This internal service charges only durably recorded work, so worker loss can leave an unbilled tail; repeated or late reports cannot silently create a second charge.
04Serve one bounded request before scaling
Start with an API gateway, a replicated control database and one inference worker. The worker loads an immutable model version that fits its accelerator memory. Durable artifact storage holds the model files and configuration needed to reload it. An accelerator such as a graphics processing unit, or GPU, performs the model calculations.
The gateway authenticates the caller, applies the model's exact prompt template, tokenizes the resulting input and validates context and output limits. It records G81 under a unique tenant/request ID. A short bounded queue feeds the worker; the first implementation handles one generation at a time. Excess work receives a retryable overload response rather than waiting indefinitely.
The worker first performs prefill: it processes the prompt and builds cached attention keys and values. During decode, it repeatedly computes the next output token using the existing prefix and newly generated tokens. The key/value, or KV, cache avoids rebuilding the whole prefix's attention state for every next token. The worker sends numbered output events through the gateway until an end condition, output limit, deadline or cancellation stops generation.
The gateway returns final status and recorded usage. If G81's worker crashes, the status becomes interrupted and its partial output remains visibly incomplete.
The control database preserves generation status; the worker owns prefill, decode and live KV memory. Artifact storage reloads a model, not an interrupted stream.
Read each connection in order
- syncPrompt and limitsClient → Generation gateway
- syncRecord identity and assignmentGeneration gateway → Generation control DB
- syncAdmit bounded generationGeneration gateway → Model worker and KV memory
- asyncLoad exact versionVersioned model artifacts → Model worker and KV memory
- returnNumbered output and usageModel worker and KV memory → Generation gateway
- returnStream and final statusGeneration gateway → Client
05Calculate token work, active streams and memory
Assume 100 requests per second with G81's 4,000 input and 600 output tokens: 400,000 input tokens and 60,000 output tokens per second. If an illustrative worker separately sustains 15,000 input tokens per second or 1,500 output tokens per second, the isolated lower bounds are 27 prefill workers and 40 decode workers. Forty mixed workers are not thereby sufficient; both phases compete for resources. Benchmark the combined workload with headroom.
At fifty output tokens per second per stream, 600 output tokens take about twelve seconds. Stable traffic therefore has roughly 100 × 12 = 1,200 active decoders, before accounting for queued or prefilling requests. Longer outputs occupy memory and execution capacity longer even at unchanged request rate.
KV-memory calculation for the sample architecture
| Input | Assumed value |
|---|---|
| Layers | 32 |
| Stored KV heads per layer | 8 |
| Elements per head | 128 |
| Bytes per element | 2 |
Bytes per token = 2 × 32 × 8 × 128 × 2
= 131,072 bytes
= 128 KiB
The first factor counts keys and values. Use the stored KV-head count, which may differ from the query-head count. A fully grown 4,600-token sequence needs about 575 MiB, excluding model weights and other execution memory.
A 20 GiB pool reserved for KV can hold only 35 such maximum-length sequences before additional overhead. Define latency targets for a tested input/output distribution, not arbitrary prompt lengths.
06Separate durable identity from temporary execution
Start a generation
POST /generations
| Request field | Meaning |
|---|---|
| Request ID | Client retry identity within the authenticated tenant. |
| Immutable model version | The exact model selected for the generation. |
| Input | The prompt to process with that model’s effective settings. |
| Maximum output tokens | A bound on how much output to generate. |
A repeated tenant/request ID with the same effective settings returns the existing generation; different content conflicts.
Inspect G81
GET /generations/G81
Return generation status after checking ownership. All stream access checks ownership too.
Request cancellation
DELETE /generations/G81
Record an idempotent cancellation request. Invalid limits, tenant quota exhaustion and unavailable worker capacity are distinct errors.
Durable control record
| Stored fields | Purpose |
|---|---|
| Generation identity and input/settings hash | Find the same logical request and reject conflicting reuse. |
| Pinned model version | Preserve the selected execution configuration. |
| Assigned worker, deadline and state | Track ownership and the generation’s lifecycle. |
| Cumulative usage | Preserve the recorded accounting total. |
Use three database replicas across failure domains and acknowledge updates after a durable majority commit. Loss of a majority stops new admissions and authoritative state changes. Model input requires an explicit private-retention policy; routine logs contain identifiers and lengths rather than raw prompts.
Temporary worker state and stream identity
| State or field | Purpose and lifetime |
|---|---|
| Tokenized input | Input representation held by the live worker. |
| KV cache and scheduler state | Current execution state; the G81 database row does not make it durable. |
| Bounded recent-event buffer | Allows reconnect replay only while the live worker retains those events. |
| Stream event generation ID and increasing sequence number | Identify the event’s generation and position in its stream. |
Expired replay buffers or worker loss mean interruption.
Persist assignment to a specific worker process incarnation: an identifier created on each process start. That incarnation deduplicates repeated dispatch; a restarted process rejects assignments naming the old incarnation. Accept output and status callbacks only from the assigned incarnation while the generation is nonterminal. Never transparently reassign a running generation; a new attempt has a new identity.
07Use admission and batching to protect latency
Time to first token, or TTFT, measures arrival to the first output token. Inter-token latency, or ITL, measures gaps between later output tokens. Suppose the interactive targets are one-second p95 TTFT and 50-millisecond p95 ITL for the tested traffic mix. Queue delay consumes TTFT; a long prefill can also interrupt output for already active requests.
First, bound tenant input/output tokens, concurrency and queue waiting time. The router can use worker capacity reports, but the worker makes the final guarded reservation against its actual memory budget. Two gateways seeing one free slot must not both consume it. Initially reserve enough KV capacity for each admitted request's maximum allowed length; this is conservative but straightforward.
Continuous batching lets the scheduler insert new sequences and remove finished ones between execution iterations. A short request need not wait for the longest member of an original fixed batch to finish. However, simply enlarging the batch can improve aggregate throughput while worsening each stream's latency.
Chunked prefill divides a long prompt's prefill computation across bounded scheduling intervals. This gives active decoders opportunities to produce output between those intervals. It does not split the user's prompt into independent questions or remove attention to the prior context. Tune the token budget against both TTFT and ITL; prioritizing ongoing decode too strongly can make new requests wait.
Use a mature serving runtime for these mechanisms. Explain the bottleneck each change addresses and measure the result. Keep long-running batch jobs in a separate queue with a different latency budget.
08Reuse computation without losing memory ownership
A paged KV allocator stores attention state in fixed-size blocks and records which blocks hold each sequence’s tokens. A growing sequence can acquire another block instead of needing one large continuous free region. This reduces wasted space, but each distinct live token still needs storage. Initially reserve for the maximum sequence length; relax that policy only after measuring memory pressure and deciding how to pause or restart work when memory runs short.
Prefix caching reuses prefill state for identical leading tokens, such as repeated instructions and the same source document. It saves input computation, not the work of generating G81's new 600-token answer. Cache identity must include compatible model weights, tokenizer, template, adapter settings and exact effective tokens. A server-controlled tenant scope prevents private prefix reuse across unrelated tenants.
Retaining a 2,000-token prefix at 128 KiB per token occupies 250 MiB. Decide whether repeated prefill savings justify that space instead of another active sequence. Eviction stops the cache from keeping a block, but running computations may still need it. Wait for them before reusing its memory. A shared prefix stays read-only: to append tokens, copy a partial block into private storage or share only completed blocks.
Cancellation follows memory ownership
- Stop future scheduling. Mark G81 unschedulable.
- Drain active use. Let any already running kernel finish using its blocks.
- Release memory. Reclaim the blocks only after that use has ended.
Freeing immediately when the HTTP connection closes can let a new request overwrite memory still being read by the old computation.
A cancellation request prevents future iterations. Memory becomes reusable only after already running execution has released its references.
Read each connection in order
- syncStart decode iteration for G81Worker scheduler → GPU execution
- syncCancel G81Client → Gateway
- syncForward recorded cancellationGateway → Worker scheduler
- syncPrevent further G81 iterationsWorker scheduler → Worker scheduler
- returnCurrent iteration finishesGPU execution → Worker scheduler
- syncRelease eligible KV blocksWorker scheduler → Worker scheduler
- returnStopped; final recorded usageWorker scheduler → Gateway
- returnTerminal canceled statusGateway → Client
09Add ready replicas before specialized execution
When one worker reaches its tested capacity, add compatible worker replicas and route independent generations among them. Readiness means the exact model and tokenizer are loaded and a representative request succeeds, not merely that the process has started. Each worker owns its scheduler and memory pool. Keep spare ready capacity for the stated worker-loss target; a cold artifact download is not immediate failover capacity.
For each model version, route using queue delay and token/memory pressure rather than open connection count alone. A worker can reject a reservation when its local capacity has changed. An explicit rejection before execution allows reassignment. A lost admission response is uncertain, so report interruption rather than risk a second execution. Running generations are never transparently reassigned.
If the model itself does not fit one device, tensor parallelism divides layer computations across devices and introduces communication between them. A coordinated group then acts as one worker, and losing a required device can interrupt that group's requests. If the model already fits, replication is generally the simpler way to increase independent-request throughput.
Quantization reduces numerical storage precision and may save weight or KV memory, but needs compatible kernels and quality evaluation. Separating prefill and decode onto different pools is a later alternative: transferring G81's 4,000-token prefix at the example KV size moves about 500 MiB per request.
The gateway records one generation identity. Bounded queues and token-aware routing feed a ready worker, which owns its scheduler and reserves its own KV memory. Output returns through the gateway. Replicas add request capacity; they do not resume a generation lost with another worker.
Read each connection in order
- syncPrompt and generation limitsAPI client → Generation gateway
- syncRecord identity and statusGeneration gateway → Identity, assignment and status
- asyncAdmit within queue budgetGeneration gateway → Bounded request queues
- asyncDispatch eligible generationBounded request queues → Token-aware admission router
- syncRecord worker assignmentToken-aware admission router → Identity, assignment and status
- syncRequest guarded reservationToken-aware admission router → Ready model worker replicas
- asyncLoad and verify exact versionVersioned model artifacts → Ready model worker replicas
- returnNumbered output and usageReady model worker replicas → Generation gateway
- returnStream and final statusGeneration gateway → API client
10Make interruption, backpressure and accounting explicit
If a worker incarnation disappears, mark all its nonterminal generations interrupted and close their streams. A restarted worker registers a new incarnation and cannot reclaim those assignments. Do not copy only a saved prompt to another worker and claim the old continuation survived. New requests can use ready replicas; a replacement attempt is visibly separate. A durable control row preserves identity, not lost GPU tensors.
A slow client gets a bounded event buffer. If it cannot catch up within the declared grace period, cancel its generation rather than accumulating unlimited output or continuing expensive work forever. A cancel response initially means the request was recorded; the terminal status becomes canceled after the worker acknowledges stopping future scheduling. Worker loss may instead produce interrupted status. Every generation also has an absolute deadline, limiting abandoned execution when control communication fails.
For usage, periodically save cumulative input/output counts and advance the durable total monotonically. Reports of 100 and then 120 tokens mean 120, not 220. On ordinary completion, persist final usage before sending the final usage event. This internal service charges only durably recorded work: a crash at token 120 with only 100 saved leaves an unbilled twenty-token tail. Finalize the interrupted total so a late report cannot silently add a charge; record late physical work separately for capacity analysis.
During a control-database outage, stop new admissions. Existing generations can run within their already granted bounds and deadlines, but status, cancellation and final accounting may be delayed. Report that limitation instead of claiming instantaneous cancellation through a partition.
11Measure useful throughput and roll out complete versions
Measure queue time, TTFT, ITL, completion latency, input/output tokens per second, active sequences, KV occupancy, prefix reuse and cancellation delay. Break results down by prompt length, output length, model version and tenant. High GPU utilization alone can hide unacceptable token gaps or repeated wasted work.
Pin the weights, tokenizer, template, adapters and compatible runtime as one tested deployment configuration. Load a new version on separate workers, run quality and mixed-load tests, route a small canary, then expand traffic. Drain old workers until their active requests finish or reach the documented deadline. Replacing weights beneath a live KV cache invalidates the state used for continuation.
Exercise worker crashes after partial output, cancellation during a kernel, simultaneous final-slot admissions, expired reconnect buffers and repeated usage reports. A basic fairness test places a short request behind a burst of long prompts and measures both its first-token delay and existing streams' token gaps.
Protect artifact integrity and private prompts. Cache tenancy comes from trusted identity, not a client-supplied value. Generated output must not execute as server code. Compare optimizations using successful workload throughput within latency and quality targets, including memory and warm-spare costs; a larger benchmark token count is not automatically a better user experience.
12Check the design against its requirements
Before closing, check the final design against the agreed requirements. FR means functional requirement and NFR means non-functional requirement; the numbers refer to the lists above. These are proposed validation checks, not test results.
| Requirement | Mechanism in the final design | Validation and remaining limit |
|---|---|---|
| FR 1, 2; NFR 3, 5 | Durable request identity, pinned versions and one worker-incarnation assignment identify each stream. | Repeat a request, conflict its settings, read another tenant’s status and crash its worker after partial output. Require one identity and explicit interruption, not a stitched replacement answer. |
| NFR 1, 2, 4 | Worker-local memory reservations, bounded queues and measured batching protect interactive service. | Benchmark the mixed token workload, first-token p95 and later-token p95 with a burst of long prompts. Isolated phase throughput and memory arithmetic do not prove mixed capacity. |
| FR 3; NFR 3, 4 | Cancellation stops future scheduling before in-flight work releases its memory. | Cancel during a kernel and lose control communication; verify no early memory reuse and the declared delayed/interruptible outcome. |
| FR 4; NFR 5 | Ready versioned replicas and draining keep rollouts compatible. | Canary the model/tokenizer/template together and fail a worker. Check quality, tenant cache isolation and warm spare capacity; cold workers are not ready replacements. |
| FR 4; NFR 6 | Monotonic cumulative usage and terminal finalization define the charged amount. | Deliver usage 100 then 120 twice, then a late report after interruption. Require the saved total, no double addition and an explicit unbilled crash tail. |
13Rapid revision
Rehearse the numbered functional requirements and non-functional targets first. Use this table to recall the mechanisms, then close with the requirements check above.
Remember: A capacity report is a hint; the worker reserves memory.
| Concept | Role in the design | Boundary to remember |
|---|---|---|
| Prefill | Process the prompt and save its attention state. | Long prompts can delay existing streams unless scheduling limits that work. |
| Decode | Generate each next token using the prompt and tokens produced so far. | Longer outputs run longer and require more KV memory. |
| TTFT and ITL | Measure delay to the first token and gaps between later tokens. | High total throughput can still leave individual streams slow. |
| Admission | Limit token work and waiting time; reserve worker memory. | The worker reserves actual capacity; a router report may be stale. |
| Continuous batching | Add new requests and remove finished ones between iterations. | Larger batches can increase throughput while slowing each stream. |
| Chunked prefill | Process a long prompt in pieces, allowing decode work between them. | Each chunk continues the same prompt computation; chunks are not independent requests. |
| Paged KV allocation | Store growing sequences in fixed-size memory blocks. | Each live token still needs attention-state memory. |
| Prefix cache | Reuse compatible prompt computation within its permitted tenant scope. | It does not cache the new answer; running requests keep shared blocks allocated. |
| Cancellation | Stop new work, wait for running computations, then free their memory. | Closing the connection alone does not stop computation. |
| Worker loss | Keep saved status and report the unfinished generation as interrupted. | A database row cannot restore lost live KV memory. |
| Usage reporting | Save increasing cumulative counts, then fix the final total. | Repeated reports do not add charges; work lost before recording is unbilled internal cost. |
| Scaling | Add workers with compatible models; split a model only when capacity requires it. | An unloaded model or an untested throughput estimate is not usable capacity. |
14Close with a measured capacity and failure contract
Trace G81 from token counting through admission, prefill, decode, numbered output and cleanup. Explain the pinned model version, worker-owned reservation and safe cancellation.
Next, mix short prompts with G81-sized requests while injecting worker failure and cancellation. Measure useful token throughput within both latency targets. Exact continuation after worker loss remains an Advanced requirement.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What are prefill and decode, and why measure them separately?
Reveal a model answer
Prefill processes input tokens and builds cached attention state. Decode repeatedly generates the next token using the prefix and cache. A long prefill can delay existing decoders, so first-token latency and gaps between output tokens reveal different scheduling problems.
Interviewer follow-up
Does prefix caching make the whole response free?
Reveal the follow-up answer
It reuses compatible input computation. The new output still requires decode work, and retaining the prefix occupies real KV memory.
What the answer must demonstrate: Define both execution phases and connect their interference to first-token and inter-token latency.
Do isolated bounds of 27 prefill workers and 40 decode workers prove forty workers are enough?
Reveal a model answer
No. The isolated measurements each assume a particular workload, while mixed workers share compute, memory bandwidth and capacity between phases. Use those figures as lower bounds, then load-test the combined prompt/output distribution with headroom and latency targets.
Interviewer follow-up
Why calculate active streams?
Reveal the follow-up answer
Each live sequence occupies KV state. At 100 arrivals per second and about twelve seconds of decode, roughly 1,200 sequences are active before queue and prefill time are included.
What the answer must demonstrate: Treat isolated throughput as lower bounds, then account for mixed execution and live-sequence memory.
Two gateways see the same worker with one free slot. What prevents over-admission?
Reveal a model answer
The worker makes a guarded reservation against its actual sequence and KV capacity before accepting execution. Registry reports guide routing but do not allocate memory. One reservation succeeds; the other must wait within its deadline or use another eligible worker before execution starts.
Interviewer follow-up
Why start with maximum-length reservation?
Reveal the follow-up answer
It provides a straightforward memory bound using the input plus allowed output length. Dynamic allocation may use memory better but needs explicit handling when sequences grow and capacity runs out.
What the answer must demonstrate: Put the guarded allocation at the worker; explain the simplicity and utilization cost of maximum-length reservation.
How do continuous batching and chunked prefill solve different problems?
Reveal a model answer
Continuous batching lets finished sequences leave and new ones enter between iterations, avoiding empty fixed-batch slots. Chunked prefill limits how much long-prompt work runs at once so active decoders get opportunities to produce output.
Interviewer follow-up
Can either optimization worsen latency?
Reveal the follow-up answer
Yes. Larger batches can lengthen token gaps, and strong decode priority can delay new prefills. Measure both first-token and inter-token latency for the actual workload.
What the answer must demonstrate: Distinguish dynamic batch membership from prefill scheduling and measure both latency consequences.
Why is it unsafe to free the KV cache as soon as a client disconnects?
Reveal a model answer
A GPU kernel may still be reading those blocks. Immediate reuse by another request can corrupt live execution. Record cancellation, stop future scheduling, wait for in-flight references to finish, and only then release blocks.
Interviewer follow-up
What if the worker is unreachable?
Reveal the follow-up answer
Record cancellation intent and rely on the bounded generation deadline while establishing worker loss. Do not claim physical work stopped instantly; status may become interrupted rather than acknowledged canceled.
What the answer must demonstrate: Preserve in-flight memory references and distinguish recorded cancellation intent from confirmed stopped execution.
Which identity is required for safe prefix reuse?
Reveal a model answer
Exact effective prefix tokens, compatible model weights/tokenizer/template/adapters, and a server-controlled tenant or approved trust scope. Shared blocks remain immutable while live requests reference them. A cache entry is computation state, not an independently authorized answer.
Interviewer follow-up
Why not let clients choose the tenancy salt?
Reveal the follow-up answer
A client could select another tenant's scope and create cross-tenant reuse or timing exposure. The gateway derives isolation settings from authenticated identity.
What the answer must demonstrate: Include exact compatible prefix identity and trusted tenant scope, with immutable shared live state.
What survives a worker crash in the selected design?
Reveal a model answer
The generation identity, status and recorded usage survive in durable storage. Live KV tensors and unsaved output events disappear. Mark the dead process’s generations interrupted. A restarted worker has a new incarnation—an identity for that process start—and rejects old assignments; a replacement generation is an explicit new attempt.
Interviewer follow-up
Why is saving only the prompt insufficient?
Reveal the follow-up answer
It can start another computation, but does not restore the exact prior execution state or guarantee the same sampled continuation.
What the answer must demonstrate: Separate durable identity from ephemeral state, bind assignments to process incarnations, and make replacement attempts explicit.
How do cumulative usage reports avoid duplicate charges?
Reveal a model answer
Advance the saved count monotonically for one generation: reports of 100 and 120 output tokens yield 120, not 220. Finalize a terminal count so late reports cannot reopen billing. Under this design, work after the last durable report at a crash is an internal unbilled cost.
Interviewer follow-up
What should a rollout verify besides throughput?
Reveal the follow-up answer
Output quality, tokenizer/template compatibility, mixed-load first-token and inter-token latency, memory, cancellation and recovery. New workers load the tested version while old workers drain.
What the answer must demonstrate: Advance cumulative counts rather than adding them, finalize terminal totals, and disclose unrecorded crash-tail work.
Blank-page exercise · 45 minutes
Build the answer yourself
Design a multi-tenant text-generation service receiving 100 requests per second, each averaging 4,000 input and 600 output tokens. Start with one bounded worker, then justify scheduling and scaling changes. Trace cancellation during execution and a worker crash after partial output.
- Agree the numbered functional requirements and non-functional targets: streaming, cancellation, token load, both latency measures, privacy and interruption behavior.
- Derive token throughput and memory budgets, then choose scheduling and scaling changes.
- Handle cancellation, worker loss and usage reporting without pretending lost execution state can be resumed.
- Calculate both token rates and architecture-dependent KV memory.
- Validate generation, retry, cancellation, rollout, latency, memory and recorded usage against the numbered requirements; identify mixed-load tests and worker-loss limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design an LLM inference platformTwo gateways see one free worker slot. Who decides which generation may start?Recall first, then reveal
The worker reserves its actual KV memory before starting work. Router reports can be stale. Admission also bounds input/output tokens and queue waiting, so request count alone is insufficient.
A capacity report is a hint; the worker reserves memory.
Return to lessonDesign an LLM inference platformThe client cancels while a GPU computation still uses its KV blocks. When can memory be freed?Recall first, then reveal
Stop new work, wait until running computations release the blocks, then free them.
Stop → drain → free.
Return to lessonDesign an LLM inference platformDoes a saved generation record let a new worker resume the exact computation?Recall first, then reveal
No. It does not contain live GPU KV memory or output events that were never saved.
Identity survives; execution may not.
Return to lessonFinal revision
Summary and interview notes
Limit token work, waiting time and memory before generation starts. Measure first-token delay and later token gaps separately, then improve scheduling or add workers where the measurements show a bottleneck.
Remember these points
- Agree the numbered functional requirements and non-functional targets before designing components; validate the final design against them.
- Budget input tokens, output tokens and KV memory instead of treating all requests alike.
- Measure first-token latency and later token gaps independently.
- Let the worker reserve actual memory, then use continuous batching and chunked prefill for measured bottlenecks.
- Scope prefix reuse to compatible versions and trusted tenants; cancellation drains execution before memory reuse.
- Persist identity and recorded usage while explicitly interrupting streams whose live execution state is lost.
Interview tips
- State the contract before choosing storage.
- Follow one concrete request through commit, response and recovery.
Important qualifications
- Traffic and latency figures are interview assumptions, not claims about a named company's deployment.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Fine-grained block sharing and cache identity
Study copy-on-write, reference races and canonical compatibility checks beyond the main allocator ownership explanation.
- Execution ownership during control partitions
Add stronger lease and fencing behavior if transparent reassignment or stricter partition guarantees are required.
- Specialized parallelism and decoding optimizations
Compare pipeline parallelism, disaggregated prefill/decode, KV quantization and speculative decoding after measuring the core service.
- Detailed accounting and finalization protocol
Extend cumulative counters into explicit per-attempt billing reservations and terminal accounting rules.
Technical references
- vLLM optimization and tuningDocuments chunked prefill, scheduling tradeoffs, preemption, and parallelism choices.
- vLLM automatic prefix cachingDocuments compatible block reuse and cache-salt isolation.
- PagedAttention paperPrimary research on efficient attention-memory management for serving.
- vLLM documentationOfficial serving capabilities, including continuous batching and streaming; settings evolve by version.
Practice marks stay in this browser.