System designby Learnastra

System-design interview · Extended interviews

Design an LLM inference platform

By Anup Rai

Design token-based admission, prefill and decode scheduling, GPU and KV-cache budgets, streaming, cancellation and tenant fairness.

You will learn to

  • Explain the separate prompt-processing and token-generation phases.
  • Calculate token throughput and attention-memory demand instead of sizing by QPS alone.
  • Design fair admission, private prefix reuse, cancellation, and model-version rollout.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Capacity estimation: throughput, latency, concurrency and storage · Message queues, event logs, delivery guarantees, and backpressure · Authentication, authorization, and tenant isolation

Workload and timing examples are interview assumptions.

01Problem and scope

An LLM inference platform serves versioned models within latency, memory and throughput budgets. Request count alone is insufficient: a 4,000-token prompt with 600 requested output tokens consumes different prefill, decode and KV-cache resources from a short interactive request. Accept a request only when its token and memory costs fit the available budget. Specify which streamed events a client can resume, and return a defined overload error when the budget is exhausted. Training, fine-tuning orchestration and execution of generated tool calls are outside this serving scope.

Define the serving metrics and state

A token is a unit produced by the model's tokenizer, often a word fragment. Prefill processes input tokens and builds attention state. Decode repeatedly produces the next output token using that state and prior output. Time to first token (TTFT) is the delay from request arrival to the first output token; inter-token latency (ITL) is the delay between successive output tokens. A KV cache stores attention keys and values so subsequent decoding does not recompute the entire prefix from scratch.

Clarify training, tools and model scope

Candidate: “Do we train models, execute generated tools, or only serve text?” Interviewer: “Serve versioned text models, stream tokens, support tenants and cancellation.” Candidate: “I will start with one worker and bounded admission, then measure mixed prefill/decode load before adding replicas or parallelism.” Training, fine-tuning orchestration and autonomous tool execution are outside scope. A model-produced tool-call proposal is output data; another authorized application decides whether to execute it.

02Functional requirements

  1. Admit a generation. Record one generation for the tenant/request identity and reserve its allowed work budget.
  2. Stream ordered events. Events have generation ID and increasing sequence; no mixed attempts.
  3. Retry a request. Return existing generation/status for same payload identity.
  4. Cancel a generation. Persist intent and remove the sequence from future scheduling promptly.
  5. Change model versions. New requests use an explicit version; active requests keep their pinned version.
  6. Account for usage. Record cumulative usage counters durably without counting repeated reports twice; state who pays for work completed after the last saved counter if the worker crashes.

Generation lifecycle

A tenant submits a request with an immutable model version, input and maximum output tokens. The service validates limits and either admits a bounded generation or rejects it before promising unlimited queued work. The client receives ordered stream events and a final reason such as completed, output-limit, canceled or interrupted. It can inspect status and cancel; cancellation stops future work at a safe execution boundary rather than undoing already delivered tokens.

Constraints and exclusions

Specify reconnect behavior. This design retains a bounded stream buffer for reconnect while its worker lives; after worker loss, the generation is interrupted unless output events were separately persisted under a stronger tier. Reusing a request ID does not make stochastic computation reproduce identical text. The service does not silently restart and concatenate a new answer after an old partial stream. Tenant quotas include input/output token budgets, concurrency and queue work, not merely request count. Slow clients cannot accumulate an unbounded per-stream buffer in server memory.

03Non-functional requirements

  1. Interactive latency. One-second p95 time to first token (TTFT) and 50 ms p95 inter-token latency (ITL) for the defined interactive mix. Also measure completion time and successful-output rate.
  2. Admission and queue bound. Targets apply to admitted requests under tested limits, not arbitrary million-token prompts. Budget queue time; reject or offer a batch tier when predicted waiting consumes most of the TTFT allowance.
  3. Availability. Target 99.9% availability for authenticated API requests within the documented request and workload limits; continue service after one worker/zone failure using reserved capacity.
  4. Control durability. Replicate request identity, terminal status and usage so they survive the declared zone failure.
  5. Execution limits. Bound context length, generated length and total active KV blocks.
  6. Regional recovery. Include compatible weights, tokenizer/runtime and warm capacity in the recovery plan; cold model reload time contributes to the recovery objective.

Execution invariants

Boundary Required guarantee
Generation ownership Only the worker assigned the current execution epoch may publish current results; the epoch is a version number that changes when ownership changes
Client stream Never merge events from different attempts into one apparent stream
Prefix reuse Never share private prefix state across incompatible model or tenant contexts
Cancellation Do not free memory still referenced by an in-flight kernel
Model quality Pin tested weight/tokenizer/template settings; stochastic output can still vary

Ephemeral execution and measured capacity

Live GPU KV state is ephemeral here: a worker crash interrupts its active streams even though control metadata survives. Hardware, precision, batch mix and model architecture determine capacity. This chapter uses illustrative measurements, not named-GPU performance claims.

04Capacity estimates

At fifty generated tokens/s per stream, 600 tokens take about twelve seconds. Stable arrivals at 100/s imply roughly 100 × 12 = 1,200 active decoders before queued/prefill requests. Longer outputs occupy slots and memory longer even if request QPS stays unchanged.

The KV calculation counts the stored attention state for each token across the model’s layers. A KV head contributes a key vector and a value vector; the head dimension is the number of elements in each vector. Multiplying those counts by bytes per element gives storage per token. Use the stored KV-head count, which need not equal the query-head count.

For an illustrative attention architecture with 32 layers, eight KV heads, head dimension 128 and two-byte elements, KV bytes per token are 2 × 32 × 8 × 128 × 2 = 131,072, or 128 KiB. The initial factor two is for keys and values. A full 4,600-token sequence needs about 575 MiB; 1,200 such fully grown sequences would need about 674 GiB. Average active length is lower, and architectures/parallelism/quantization change the allocation. Weights, activations, communication buffers and runtime reserve are additional.

Resource illustration Result Admission consequence
20 GiB usable KV pool 20 GiB / 575 MiB ≈ 35 maximum-length sequences Cannot admit unlimited concurrent contexts
2,000-token prefix at 128 KiB/token 250 MiB cached state Reuse saves compute but occupies real memory
100 canceled streams, 575 MiB each Up to about 56 GiB released eventually Cancellation latency affects useful capacity

These estimates explain token and block budgets before choosing a serving framework.

05APIs and contracts

Effective input and request identity

Request A sends POST /v1/generations with {"requestId":"generation-81","model":"summarizer-v4","input":"...","maxOutputTokens":600,"stream":true}. Tenant identity comes from authentication. The gateway applies the exact versioned prompt template and tokenizer before computing limits; counting only the visible user text would miss system/tool-format overhead. A reused request ID with changed effective input/settings returns 409.

Generation interfaces

Interface Contract
Generation response generationId:g81, queued/running state and ordered token events
Stream event {generationId:g81,epoch:7,sequence:121,text:"..."}
Final event Terminal reason, durably recorded billable usage and pinned model/template versions
DELETE /v1/generations/g81 Idempotent cancellation intent; terminal state remains inspectable
GET /v1/generations/g81 Current state and declared reconnect/interruption behavior
GET /v1/models Authorized immutable versions and supported limits

Errors, resume and usage semantics

Reject invalid settings/context with 400, tenant exhaustion with 429 and unavailable compatible capacity with 503. Retry hints include jitter expectations, and batch requests may use a different latency tier. A bounded retained stream buffer can resume from an event sequence only while those events remain available; if the cursor expired, report it. Do not promise replay of unpersisted output after a worker crash. Usage semantics must distinguish processed input, generated output, delivered output and cached-input discounts if any; these are product accounting choices, not inferred from packet count.

Durable-watermark billing

A usage watermark is a saved cumulative count, such as 100 output tokens processed so far. Later reports advance that count; receiving the same report twice must not double the charge. The crash tail is work completed after the last saved report and lost from accounting when a worker fails.

  • Define billable work. For this design, billing uses only durably recorded cumulative work watermarks, not all physically executed work.
  • Advance durable counters. Each attempt periodically reports cumulative input/output work; the accounting authority advances each counter monotonically for its (generationId, epoch) and deduplicates event identity.
  • Flush final usage. On normal completion or cancellation, flush and acknowledge the final watermark before sending a final usage event.
  • Account for the crash tail. If the worker crashes after token 120 while only 100 are durably recorded, the unreported 20-token tail is an unbilled internal cost.
  • Disclose interrupted usage. The interrupted result labels its recorded usage accordingly.
  • State the throughput/accounting tradeoff. This avoids a durable round trip per token; it accepts some underbilling rather than claiming exact crash-proof accounting from an asynchronous stream.

Seal terminal accounting

  • Seal at terminalization. When the attempt completes, is canceled or is declared interrupted, the accounting transaction marks its billable totals final. This is called sealing the attempt.
  • Record final totals and release reservation. It atomically records the final billable watermarks and releases unused reservation; later stale reports cannot reopen or increase that sealed invoice.
  • Accept pre-seal reports. Before sealing, a delayed authenticated report can advance a watermark monotonically.
  • Reconcile post-seal work internally. After an interrupted attempt is sealed, a late usage report contributes only to the operator’s estimate of actual compute consumed; it cannot add a new user charge.
  • Keep the crash-tail policy stable. This makes the stated unbilled crash-tail policy stable across delayed messages.
  • Separate physical recompute from logical billing. Here internal preemption/recompute and retransmitted stream events are not new billable logical tokens; report that physical work separately for capacity analysis.

06Data model and access patterns

Durable entities and registry

  • Generation identity. Generation(tenantId,requestId,generationId,payloadHash,modelVersion,state,executionEpoch,cancelRequested,budgetReservation) owns request identity.
  • Execution attempt. Attempt(generationId,epoch,workerId,startedAt,terminalReason) identifies execution.
  • Usage identity. UsageEvent(generationId,epoch,eventSequence,cumulativeInput,cumulativeOutput) has a unique identity.
  • Accounting rule. The accounting authority takes monotonic per-attempt watermarks, rather than adding cumulative counters as if they were independent deltas.
  • Compatible model version. ModelVersion records weight checksum, tokenizer, template, adapter and runtime compatibility.
  • Worker registry. A worker registry advertises loaded compatible versions, health and approximate available token/block capacity.

Ephemeral worker state and epochs

A KV block is a fixed-capacity allocation for cached key/value elements. A block table tells the worker which physical blocks hold a sequence’s logical token positions. This separation lets a growing sequence use available blocks without requiring one large contiguous allocation; the worker must still track every live reference before reusing a block.

Worker-local state contains tokenized input, active sequence positions, KV block tables, scheduler queues and bounded stream buffers. This state is not made durable just by saving a Generation row. On worker failure, the control plane marks the attempt interrupted and either leaves retry to the client or creates an explicitly separate attempt under a documented policy. A time-limited ownership lease and its epoch number let the control store reject terminal-state updates from a replaced worker; the stream gateway also rejects output carrying an old epoch.

Prefix identity and model artifacts

A prefix is the leading token sequence shared by two requests, such as the same instructions and source document. Prefix caching retains the prefill state computed for those tokens so a compatible request can start from it. It reuses prior input computation; it does not supply the new request’s generated answer.

Prefix cache keys incorporate exact effective tokens, compatible model/weight/tokenizer/adapter settings and a server-controlled tenant or trust-group scope. A prefix cache reuses internal tensors; an answer cache returns text and requires a different correctness policy. Raw private prompts and tensors are not shared through an unscoped key. Durable weights live in an artifact store with checksums and access control; loading a similarly named model without checking version compatibility would break both output consistency and cache safety.

07Basic working design

One model and one worker

Start with one authenticated API and one inference worker hosting a model that fits its hardware. Tokenize request A's effective prompt, validate the 4,000-plus-600 token bound, and admit only if the queue and memory budget can support it. The worker performs prefill, then decodes tokens one iteration at a time and streams numbered events. Request B waits behind request A in a simple first-in queue; that is inefficient but understandable.

Lifecycle, cancellation and worker loss

The API records g81 before scheduling and returns a terminal outcome only when completion/cancellation/interruption is known. A client disconnect triggers cancellation under a stated grace policy. The worker releases KV memory after it is no longer used by execution, and the accounting path bills durably recorded logical token work rather than maximum reserved tokens; unrecorded crash-tail work remains an internal cost. If the worker crashes, g81 is interrupted; the service does not claim its saved metadata can recreate the lost KV cache.

Measure before adding complexity

For a small internal service this can be the right starting point. Load-test prompt/output distributions, measure TTFT and ITL separately, and retain a small set of known quality prompts. A faster throughput benchmark that allows multi-second token gaps does not validate the interactive target. The baseline gives us the data needed to justify batching, prefix reuse or multiple workers, instead of guessing capacity from the GPU's memory size alone.

architecture · baselineOne admitted generation on one worker

The baseline measures real prefill/decode behavior, streams tokens and reports interruption honestly if its worker fails.

One admitted generation on one workerThe baseline measures real prefill/decode behavior, streams tokens and reports interruption honestly if its worker fails. client to api: 1. Generate / cancel; api to control: 2. Reserve request identity; api to worker: 3. Admit bounded token work; weights to worker: 4. Load compatible model; worker to client: 5. Stream ordered tokens1. Generate / cancel2. Reserve request identity3. Admit bounded token work4. Load compatible model5. Stream ordered tokensACTORRequest A andrequest B clientsSERVICEAuthenticatedgeneration APISTOREGeneration controlrecordsSERVICESingle inferenceworker and KVmemorySTOREVersioned modelartifactssynccontrol
Read each connection in order
  1. sync1. Generate / cancelRequest A and request B clients → Authenticated generation API
  2. sync2. Reserve request identityAuthenticated generation API → Generation control records
  3. sync3. Admit bounded token workAuthenticated generation API → Single inference worker and KV memory
  4. control4. Load compatible modelVersioned model artifacts → Single inference worker and KV memory
  5. sync5. Stream ordered tokensSingle inference worker and KV memory → Request A and request B clients

08Find the baseline flaws

Failure test What breaks and what must follow
Long prefill blocks short requests Request A's 4,000-token prefill monopolizes the worker while request B's short question waits. A fixed batch can create another inefficiency: all requests start together, but short completions leave empty slots until the longest finishes if the scheduler cannot add new work. The problem is scheduling at the token-iteration level, not simply insufficient HTTP threads.
KV allocation and utilization Suppose the worker has 20 GiB available for KV after weights/reserve. Admitting 100 sequences that can grow to 575 MiB requires about 56 GiB, exceeding the pool even though all input requests fit in CPU memory. “There are only 100 requests” is not a memory estimate. Paged allocation reduces waste but cannot make those live tokens free. The service needs an explicit strategy: conservative reservation, preemption/recomputation, bounded swapping or rejection, with corresponding latency consequences.
Prefix privacy and abandoned computation A third counterexample is cross-tenant prefix reuse. If a private document's cached prefix makes a guessed request noticeably faster for another tenant, timing can reveal information about cache residency. Raw tensor bytes need not be returned for a side channel to matter. Cache scope must follow trusted identity, not a client-chosen salt that can impersonate another tenant. Finally, canceling only the HTTP connection leaves abandoned generation consuming GPU blocks unless cancellation reaches the execution scheduler.

09Improve the design, step by step

1. Add token-based admission and fair bounded queues

  • Trigger: Long-prompt overload triggers quotas on input/output work, context length and concurrent reserved blocks.
  • Mechanism: Weighted tenant scheduling and queue deadlines protect interactive users.
  • Benefit, cost and alternative: This improves predictable TTFT and isolation but rejects some requests that an unbounded queue would accept and later time out. Separate batch queues are preferable for workloads that can trade delay for utilization; request-QPS-only limits remain insufficient.

2. Use continuous batching with chunked prefill

  • Trigger: Completed requests leave unused slots in a fixed batch, while a long prompt can delay tokens for existing streams.
  • Mechanism: Between iterations, replace completed requests with new ones and process long prompts in bounded pieces. Existing streams keep opportunities to decode while spare capacity handles new inputs.
  • Benefit, cost and alternative: Benefits are higher utilization and smoother output; costs are scheduling overhead, tuning and possible slower TTFT for long prompts. A simpler fixed batch suits offline homogeneous jobs. vLLM documents these mechanisms, but configuration must match the tested model/workload.

3. Manage KV memory in blocks and reuse authorized prefixes

  • Trigger: Fragmentation and repeated common prompts trigger paged allocation plus compatible prefix caching.
  • Mechanism: As sequences grow, the worker allocates blocks and counts which active sequences still reference each block. Matching prefixes within an authorized cache scope avoid repeated prefill.
  • Benefit, cost and alternative: Benefits are less wasted memory and input computation. Costs include metadata, eviction, recomputation and security scope. Full worst-case reservation is simpler but may waste capacity; unscoped reuse is rejected because it crosses privacy boundaries.

4. Add compatible replicas and controlled model parallelism

  • Trigger: Aggregate demand or model size triggers more workers.
  • Mechanism and tradeoff: Replicas scale independent requests when the model fits; tensor parallelism splits layer computation across devices when needed, adding communication. Separate prefill/decode pools are a later measured alternative, with large KV transfers and new failure modes. Use separate pools only when measured benefits justify that transfer cost. Every rollout pins immutable versions and drains active streams before retiring old workers.

Each step is judged against both useful token throughput and latency, not GPU utilization alone.

10Detailed architecture

Admission and compatible routing

The gateway authenticates, applies the pinned template/tokenizer, validates budgets and records generation identity in a replicated control store. Admission selects an allowed latency tier and reserves tenant work. The router selects a worker with the right model using readiness and estimated capacity. The worker must then reserve actual KV blocks: its registry report may already be out of date.

Worker-owned scheduling and memory

Each inference worker owns its scheduler, active sequences, KV block manager, isolated prefix cache and stream output. A worker may be one device or a coordinated group running parts of the same model. State whether losing one device interrupts the whole group, and count memory across that group. Weights/tokenizer artifacts are checksum-verified before readiness. The diagram keeps prefix cache inside this execution boundary rather than presenting it as a generic shared answer cache.

Stream and accounting boundaries

A stream gateway forwards only events with the generation's current epoch and bounds slow-consumer buffers. It propagates cancellation and disconnect policy to the owning scheduler. Usage events flow asynchronously to an idempotent accounting aggregator; critical state transitions update the control authority. Metrics report queueing, TTFT, ITL, memory and fairness independently.

Independent scaling and rollout

The main synchronous path is admission through token production, while model deployment and usage aggregation are background work. Control-store replication protects request/status identity but does not checkpoint GPU tensors. A stale worker must not publish current terminal state, and the scheduler must prevent canceled work from starting another iteration before returning its blocks to the reuse pool.

architecture · finalWork-based admission and isolated execution state

The worker owns actual KV allocation. Durable control state is distinct from ephemeral tensors and stream buffers.

Work-based admission and isolated execution stateThe worker owns actual KV allocation. Durable control state is distinct from ephemeral tensors and stream buffers. client to api: 1. Submit generation-81 / cancel g81; api to control: 2. Reserve / cancel generation under epoch; api to admit: 3. Admit token/memory budget; admit to router: 4. Schedule compatible work; router to worker: 5. Assign work; check local capacity; worker to control: 5b. Claim current generation epoch; worker to kv: 6. Reserve / reuse isolated blocks; worker to stream: 7. Emit g81 epoch 7 events; stream to control: 7b. Check current generation epoch; stream to client: 8. Stream bounded ordered output; api to worker: 9. Propagate current cancellation; worker to control: 10. Guard terminal state by current epoch; artifacts to worker: 11. Load checked model version; deploy to router: 12. Route only ready versions; worker to usage: 13. Save cumulative token-usage counts; usage to control: 14. Finalize budget usage; worker to metrics: 15. Report TTFT, ITL and blocks1. Submit generation-81 /cancel g812. Reserve / cancel generationunder epoch3. Admit token/memory budget4. Schedule compatible work5. Assign work; check localcapacity5b. Claim current generationepoch6. Reserve / reuse isolatedblocks7. Emit g81 epoch 7 events7b. Check current generationepoch8. Stream bounded orderedoutput9. Propagate currentcancellation10. Guard terminal state bycurrent epoch11. Load checked modelversion12. Route only ready versions13. Save cumulativetoken-usage counts14. Finalize budget usage15. Report TTFT, ITL and blocksACTORTenant clientsSERVICEAuthenticated APIand tokenizerG1STOREReplicatedgeneration / epochauthorityG1QUEUEToken admission andfair queuesG1SERVICECompatible-modelworker routerG1SERVICEInference schedulerand executionG2CACHEKV block managerand scoped prefixcacheG2SERVICEEpoch-checkedstream gatewayG2STOREImmutable weightsand tokenizer storeG3SERVICEModel readiness androllout controlG3WORKERIdempotent usageaggregationG3STORELatency, memory andfairness metricsG3syncasynccontrolG1 Identity, admission and durable stateG2 Model-compatible execution boundaryG3 Artifacts, rollout and accounting
Read each connection in order
  1. sync1. Submit generation-81 / cancel g81Tenant clients → Authenticated API and tokenizer
  2. sync2. Reserve / cancel generation under epochAuthenticated API and tokenizer → Replicated generation / epoch authority
  3. sync3. Admit token/memory budgetAuthenticated API and tokenizer → Token admission and fair queues
  4. async4. Schedule compatible workToken admission and fair queues → Compatible-model worker router
  5. sync5. Assign work; check local capacityCompatible-model worker router → Inference scheduler and execution
  6. sync5b. Claim current generation epochInference scheduler and execution → Replicated generation / epoch authority
  7. sync6. Reserve / reuse isolated blocksInference scheduler and execution → KV block manager and scoped prefix cache
  8. sync7. Emit g81 epoch 7 eventsInference scheduler and execution → Epoch-checked stream gateway
  9. sync7b. Check current generation epochEpoch-checked stream gateway → Replicated generation / epoch authority
  10. sync8. Stream bounded ordered outputEpoch-checked stream gateway → Tenant clients
  11. control9. Propagate current cancellationAuthenticated API and tokenizer → Inference scheduler and execution
  12. sync10. Guard terminal state by current epochInference scheduler and execution → Replicated generation / epoch authority
  13. control11. Load checked model versionImmutable weights and tokenizer store → Inference scheduler and execution
  14. control12. Route only ready versionsModel readiness and rollout control → Compatible-model worker router
  15. async13. Save cumulative token-usage countsInference scheduler and execution → Idempotent usage aggregation
  16. async14. Finalize budget usageIdempotent usage aggregation → Replicated generation / epoch authority
  17. async15. Report TTFT, ITL and blocksInference scheduler and execution → Latency, memory and fairness metrics

11Write path and acknowledgement

Before scheduling, reserve the allowed prompt/output work and memory and fix the model version. Bill only from usage counts saved under the declared durable-watermark policy.

Numbered admission and generation trace

  1. Tokenize the exact effective request. Request A authenticates under tenant t9. The API applies summarizer-v4's exact template and tokenizer, measures 4,000 input tokens, validates maxOutputTokens 600, and hashes the effective request settings.
  2. Reserve one logical generation. Atomically reserve generation-81 as g81 with its tenant budget and execution policy. A duplicate identity returns the same g81; a different payload conflicts. Admission checks queue deadline and predicted token/memory demand.
  3. Reserve worker-local capacity. Route to a ready compatible worker. The worker atomically reserves its local sequence/block budget before acknowledging execution epoch 7, so simultaneous gateway decisions cannot overcommit the same remaining slots.
  4. Reuse only compatible scoped prefixes. Look up a compatible prefix under t9's server-controlled cache scope. Reuse only valid blocks and increment references; otherwise schedule prefill. A cache hit changes work, not the model version or allowed output limit.
  5. Interleave prefill and decode. Process request A's prefill in bounded chunks interleaved with decoding for existing requests. Request B's short request can enter later iterations instead of waiting for a whole fixed batch to finish.
  6. Stream ordered output and durable usage. Decode outputs, assign stream sequence numbers and send events tagged g81/epoch 7. Stop at model end, output limit, deadline or cancellation. Periodically report cumulative work watermarks with unique usage-event identity; the durable authority advances counters monotonically.
  7. Drain execution and finalize accounting. On terminal state, stop scheduling, wait for in-flight execution to release references, free/reuse eligible blocks, flush the final usage watermark, and finalize the durable result. A normal final usage event waits for that durable acknowledgement; a crash instead reports the last recorded watermark as interrupted usage. Unused budget reservation is released under the accounting policy.

Registry estimate versus atomic reservation

Two routers may both see the same free memory. The worker’s atomic reservation lets only one claim that remaining capacity.

12Read and delivery path

Streaming obeys backpressure and disconnect policy. Cancellation stops new scheduling before freeing state still used by in-flight GPU work.

Numbered stream and cancellation flow

  1. Authorize ordered stream delivery. The client of request A subscribes to g81 and receives ordered text events. The stream gateway verifies tenant ownership and execution epoch before forwarding them. Client rendering handles event repetition by sequence where reconnect buffering permits it.
  2. Resume only retained events. The client can inspect status without creating another generation. If it reconnects within retained buffer limits, it asks after its last event sequence; otherwise it receives an explicit expired/interrupted outcome.
  3. Observe interactive fairness. The client of request B measures TTFT separately from ITL. The scheduler's fairness policy should keep the short interactive request from waiting behind an unbounded queue of long-document requests.
  4. Persist cancellation intent. The client cancels request A after token 120. The API durably marks cancelRequested and notifies epoch 7's worker. A repeated cancel is harmless; canceling an already completed request reports its terminal state rather than pretending output was undone.
  5. Stop scheduling before freeing memory. The scheduler observes cancellation at its next safe boundary, prevents new decode/prefill work for g81 and marks its stream canceled. It waits until in-flight kernels no longer reference blocks before freeing them. Some already-computed events may have been in transit; the client knows the cancellation boundary is not retroactive erasure.
  6. Finalize usage and unused reservation. Actual usage is finalized according to the declared policy, and reserved-but-unused work is released. A bounded slow-consumer policy can similarly pause briefly or cancel instead of allowing unlimited stream-buffer growth.

Worker-loss interruption contract

If a worker dies after token 120, the service reports interrupted. A new model attempt might produce a different continuation even with the same high-level question, so it is not silently appended under g81's old event sequence. Durable output replay or exact continuation would require additional checkpoint/state guarantees beyond this design.

13Correctness deep dive

Paged KV ownership

Paged KV allocation manages fixed-size blocks with ownership/reference counts. It avoids reserving one contiguous region for the maximum possible sequence. Smaller blocks reduce unusable gaps between allocations, called external fragmentation; unused slots inside a partially filled final block are internal fragmentation. It does not reduce the number of logical attention values required for distinct live tokens. Prefix reuse lets compatible requests refer to already computed blocks, while later divergent tokens allocate separate blocks.

Concept in focusShare the prefix; separate the continuations

Arrows from two requests converge on the same prefix blocks, then lead to different suffix blocks.

Share the prefix; separate the continuationsArrows from two requests converge on the same prefix blocks, then lead to different suffix blocks. Trace shared and request-specific KV block ownership. Requests A and B reference compatible prefix blocks P1 and P2. Each owns different suffix blocks; reuse and release must respect isolation and in-flight GPU work.Two requests share prefix KV blocks, then divergeRequest ARequest BP1P2A suffixB suffixShared prefixReuse requires matching model, tokens, runtime and isolation context.Recycle a block only after all owners and in-flight work release it.

Remember: Same prefix can share memory; different continuations need their own state.

Read the diagram
  1. Trace shared and request-specific KV block ownership.
  2. Requests A and B reference compatible prefix blocks P1 and P2.
  3. Each owns different suffix blocks; reuse and release must respect isolation and in-flight GPU work.
Try from memoryCan request A free prefix block P1 as soon as A finishes?

Not if B or in-flight work still uses it. Shared ownership must be accounted for before recycling the block.

Memory transition table

Event Required enforcement Result
Admit g81 Worker scheduler reserves within block/token limits No double admission of the same free capacity
Match prefix Exact compatible key and trusted tenant scope Increment references to reusable blocks
Cancel g81 Mark sequence unschedulable for current epoch No future iterations are added
Kernel still in flight Keep references until execution completion Blocks cannot be reused prematurely
Release last reference Block manager observes zero live references Return block to pool or permitted prefix cache

Cancellation race

Prefix security boundary

Prefill savings and memory cost

Prefix reuse primarily saves prefill; six hundred new output tokens still require decode work. An answer cache is separate and must include task permissions, source freshness and generation settings. Neither cache should be described as a proof of deterministic output or universal protection from every hardware side channel.

Shared-prefix copy-on-write

Shared prefix blocks are read-only while referenced by several sequences. A sequence that must append into a shared partial block first allocates and copies a private block, or the implementation shares only complete immutable blocks. It must not write new KV entries into another sequence's shared state. The block manager orders reference updates, eviction and allocation so they cannot race. Eviction releases the cache’s claim, but a block remains allocated while a computation still uses it.

Canonical cache identity

Cache identity uses a collision-resistant hash over canonical model/tokens/scope data, with validated compatibility metadata. A fast unverified hash collision must not substitute another prompt's tensors. Different attention layouts, quantization formats or adapters can change compatibility even if displayed model names match. Treat a serving framework's cache-salt and hash options as version-tested configuration, not a claim that default settings meet every tenant boundary.

sequence · cancel-safe-releaseCancellation stops scheduling before blocks are freed

A block still referenced by an in-flight kernel cannot be reused for another sequence, even after the client cancels.

Cancellation stops scheduling before blocks are freedA block still referenced by an in-flight kernel cannot be reused for another sequence, even after the client cancels. client to api: Cancel g81 after token 120; api to api: Persist cancelRequested for epoch 7; api to sched: Cancel current g81 epoch 7; sched to sched: Mark sequence unschedulable; sched to gpu: Wait for in-flight iteration boundary; gpu to sched: No active references for g81; sched to blocks: Release g81 block references; blocks to sched: Reuse only zero-reference blocks; sched to api: Save final usage; finalize canceled result; api to client: Canceled; prior output retainedPARTICIPANTRequest A clientPARTICIPANTAPI authorityPARTICIPANTWorker schedulerPARTICIPANTExecution enginePARTICIPANTKV block manager1. Cancel g81 after token 1202. PersistcancelRequested forepoch 73. Cancel current g81 epoch74. Mark sequenceunschedulable5. Wait for in-flight iterationboundary6. No active references forg817. Release g81 block references8. Reuse only zero-reference blocks9. Save final usage; finalizecanceled result10. Canceled; prior outputretainedsyncreturn
Read each connection in order
  1. syncCancel g81 after token 120Request A client → API authority
  2. syncPersist cancelRequested for epoch 7API authority → API authority
  3. syncCancel current g81 epoch 7API authority → Worker scheduler
  4. syncMark sequence unschedulableWorker scheduler → Worker scheduler
  5. syncWait for in-flight iteration boundaryWorker scheduler → Execution engine
  6. returnNo active references for g81Execution engine → Worker scheduler
  7. syncRelease g81 block referencesWorker scheduler → KV block manager
  8. returnReuse only zero-reference blocksKV block manager → Worker scheduler
  9. syncSave final usage; finalize canceled resultWorker scheduler → API authority
  10. returnCanceled; prior output retainedAPI authority → Request A client

14Failure and recovery

Failure or condition Surviving state, response and recovery
Worker crash Worker crash: Active KV state and unpersisted stream buffers disappear. The control plane marks epoch 7 interrupted after its lease/health failure is established. The client of request A retains whatever text it already received and can start an explicit new attempt. Model weights reload from durable artifacts; warm compatible replicas absorb new work within reserved capacity. The durable request and previously recorded usage survive. Physical work after the last usage watermark can be lost from accounting; under our declared policy that crash tail is unbilled, not fabricated as an exact count.
Control-authority partition Control authority partition: A minority cannot create new generation identities or safely change ownership. Existing admitted workers may continue within their bounded lease/policy, but terminal state and cancellation propagation require reconciliation. Epoch checks prevent a stale worker and replacement from both presenting one continuous current stream. In this design, workers stop scheduling new iterations and stream gateways stop admitting further events when their bounded execution/forwarding lease expires. Renewal uses the authority; local timeout checks use a conservative deadline accounting for elapsed request time and clock uncertainty. A replacement is activated only after the old forwarding/worker lease interval is fenced. Already admitted kernel work or network bytes may finish; the service does not claim instantaneous physical cancellation.
Long-prompt flood Long-prompt flood: Token budgets and tenant concurrency limits reject work before memory collapse. Weighted queues reserve interactive capacity; batch work can wait longer. A scheduler may preempt and recompute lower-priority sequences under an explicit policy, trading latency for memory. Repeated preemption indicates over-admission and should not become invisible “free” capacity.
Slow or disconnected client Slow consumer or disconnected client: Bounded buffers and cancellation free execution resources after a grace period. A gateway that drops only the socket but leaves the worker running wastes expensive tokens and blocks. Apply backoff/jitter to retries so an overloaded model is not hit by synchronized repeated prefills. A full-region outage requires capacity elsewhere with the compatible model loaded. Recovering request metadata alone does not load the weights or make another accelerator ready.

15Operations, security, and cost

Latency, token and memory metrics

Observe input and output tokens/s, queue delay, TTFT, ITL, completion latency, active sequences, occupied/free KV blocks, prefix-hit tokens, preemptions, cancellation lag and per-tenant service share. Measure useful successful tasks alongside raw tokens and hardware utilization. A worker at 99% utilization may produce unacceptable token gaps or spend much of its time recomputing preempted prefixes.

Measured capacity and cost

Cost comparisons use resource units rather than guessed device prices. If a repeated private 2,000-token prefix saves 2,000 prefill tokens, ten reuses avoid 20,000 input-token computations while retaining about 250 MiB in the illustrative architecture. Compare the worker time saved with the memory no longer available to other sequences, and measure how often the prefix is evicted. For disaggregated prefill/decode, moving a 4,000-token KV prefix at 128 KiB/token transfers about 500 MiB per request; at 100 requests/s that is roughly 49 GiB/s before transport overhead. This quantifies the network bandwidth needed between the prefill and decode pools and helps decide where to place them.

Model and tenant security

Protect model/artifact integrity, tenant cache scopes and prompt/output retention. Keep credentials out of prompts and do not let generated text become server code. Tool execution belongs to a separate authorized service. Logs should prefer IDs, lengths and error classes over raw private prompts unless a reviewed debugging policy permits content access.

Version rollout and draining

Roll out immutable weight/tokenizer/template/runtime combinations through quality regression, mixed-load latency tests, a canary and controlled routing. Drain old workers while active requests finish; do not swap weights beneath live KV state. Test cancel-during-kernel, stale epochs, model reload failure, queue saturation and stream reconnect. Verify usage-event deduplication after crashes so repeated reporting does not distort tenant budgets.

Watermarks and billing reconciliation

Track the gap between worker-reported work and durable watermarks, reporting delay and unbilled interrupted tails separately from usage-event duplication. A 100-token watermark followed by a crash at token 120 is a recovery test: charge 100 recorded output tokens, mark the result interrupted, and never add a late duplicate report twice. Monitor actual hardware work separately. Preventing duplicate billing records does not prove that every computed token was recorded.

16Decision ledger and limitations

Decision table

Decision Benefit Cost / consequence Change trigger
Continuous mixed batching Reuses slots as requests arrive/finish Scheduler complexity and contention Homogeneous offline jobs favor simpler batches
Chunked prefill More decode opportunities during long prompts Prefill scheduling overhead and tuning Tight long-prompt TTFT changes the balance
Paged KV allocation Less fragmentation and flexible growth Block metadata and lifecycle correctness Simpler fixed workload may tolerate reservation
Tenant-scoped prefix reuse Saves repeated input work Memory occupancy and isolation policy Low reuse favors earlier eviction
Early work-based rejection Predictable admitted latency Explicit client errors during peaks Batch tier can accept longer deadlines
Replicas before phase separation Simple failure boundaries and no KV network handoff Mixed-resource interference Benchmarks justify disaggregated prefill/decode

Parallelism and disaggregated serving

Tensor parallelism splits model-layer operations across devices and adds communication; pipeline parallelism places successive layers on different devices, which can sit idle while waiting for an earlier stage to produce input. Replicating complete workers is generally simpler when a model already fits and only aggregate throughput is lacking. The correct mix depends on actual weights, memory, network and latency targets, not a universal rule that one strategy is fastest.

Reservation versus utilization

Full maximum-length reservation gives a simple memory bound but may waste unused output capacity. Allocating memory as sequences grow can use space better, but may require pausing and recomputing work or tighter admission limits. It must still prevent uncontrolled out-of-memory failures. Prefix caching and answer caching solve different problems. A high input-cache hit rate does not mean output generation is cheap, and a larger batch can improve throughput while worsening each stream's latency. Report both before claiming an optimization succeeded.

Quantization and speculative decoding

Two further optimizations are worth discussing after the baseline is measured. KV quantization can reduce bytes per cached token but changes numerical behavior and needs compatible kernels, quality tests and scale/metadata accounting; the worked 128 KiB/token calculation deliberately assumes two-byte elements. Speculative decoding uses a cheaper draft process to propose multiple tokens and a target-model verification step to accept/correct them. Exact sampling preservation requires the algorithm's target verification and acceptance rules; blindly accepting draft tokens changes the model distribution. Its benefit depends on draft acceptance, verification cost and traffic mix. Neither optimization removes the tenant, budget or cancellation boundaries, and support varies with the chosen model/runtime.

17Interview closing

Rehearse the architecture and contract

“I designed a versioned multi-tenant text-generation service. A four-thousand-token prompt with six hundred output tokens consumes much more capacity than a short exchange, so I budget input tokens, output tokens and KV memory rather than only requests per second. Our example needs 400,000 input and 60,000 output tokens per second, with about twelve hundred active decoders; isolated throughput bounds are not a mixed-capacity proof.

Defend the critical boundary

“I begin with one bounded worker, then add token-based admission, continuous batching, chunked prefill and paged KV allocation. Prefix reuse is compatible-model and tenant scoped. The worker atomically owns its memory budget, and cancellation reaches the scheduler before blocks are safely released. Compatible replicas scale the service; splitting prefill and decode waits for evidence because KV transfers are large.

State the cost and next measurement

“I persist request and usage identity but explicitly mark active streams interrupted when ephemeral execution state is lost. I do not silently concatenate a different regenerated answer. My next measurement is the mixed-workload latency/memory curve and cancellation recovery under peak load.”

Answer the follow-up

If the interviewer asks for an offline bulk tier, allow longer queues and larger batches under separate capacity/budgets while protecting interactive reservations. The service objective changes; the same scheduler settings should not be assumed optimal for both tiers.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Explain prefill and decode to an interviewer.

Reveal a model answer

Prefill processes the input and constructs attention state. Decode generates the next tokens iteratively using that state. A long prompt stresses input processing; a long answer keeps generation and memory active for longer.

What the answer must demonstrate: Do not collapse every latency into one average.

Applied · Question 2

Can you size this service from 100 QPS?

Reveal a model answer

Not alone. With 4,000 input and 600 output tokens per request, it needs 400,000 input and 60,000 output tokens per second. I also estimate active sequences and KV memory, then benchmark the actual mix under latency targets.

What the answer must demonstrate: Do not treat separately measured prefill and decode throughput as capacity simultaneously available on the same worker.

Applied · Question 3

Why use continuous batching instead of waiting for a fixed batch to finish?

Reveal a model answer

Requests have different output lengths. Continuous batching removes finished sequences and admits new work between iterations, reducing idle capacity. The scheduler still limits tokens and memory so larger batches do not ruin streaming latency.

What the answer must demonstrate: Explain who waits and why.

Foundation · Question 4

Is prefix caching the same as answer caching?

Reveal a model answer

No. Prefix caching reuses compatible internal prompt state and then generates a new continuation. Answer caching returns an existing result and needs additional freshness, permission, and semantic rules.

What the answer must demonstrate: Treat cache isolation as part of authorization design.

Follow-up · Question 5

The user closes the tab at token 120. What should happen?

Reveal a model answer

The gateway propagates cancellation to the scheduler, which stops further generation and frees the sequence’s resources when safe. Stream buffers are bounded, and actual usage is recorded under the declared contract.

What the answer must demonstrate: Send cancellation to the worker scheduler and verify that it stops work and releases unused memory.

Follow-up · Question 6

A GPU worker dies halfway through the answer. Can you transparently continue on another worker?

Reveal a model answer

Not without a defined recoverable state/output protocol. Normally I mark the stream interrupted; a fresh attempt may generate different text. I must not append unrelated regenerated text to the old stream silently.

What the answer must demonstrate: State the recoverability limit of live KV state.

Applied · Question 7

Why do 27 prefill workers and 40 decode workers not prove that 40 mixed workers suffice?

Reveal a model answer

Those are lower bounds from isolated benchmarks. Prefill and decode share compute, bandwidth and KV capacity on the same workers, and the batch mix changes latency. I use them to reject obviously undersized plans, then measure representative mixed traffic with headroom under TTFT and ITL targets.

What the answer must demonstrate: Do not turn isolated maximum throughput into simultaneous guaranteed capacity.

Follow-up · Question 8

Why not free request A’s KV blocks as soon as the API receives cancel?

Reveal a model answer

A running kernel may still read those blocks. The API saves and forwards cancellation; the worker stops scheduling new work, waits for the running computation to finish and releases its references. Reusing memory sooner could corrupt another request or expose data.

What the answer must demonstrate: Cancellation acknowledgement, scheduler stop and memory reclamation are distinct moments.

Blank-page exercise · 45 minutes

Build the answer yourself

Design a multi-tenant inference service for request A’s long summary and request B’s short question, then cancel request A and lose a worker mid-stream.

  • Define prefill, decode, TTFT, ITL, and KV memory.
  • Calculate input/output token rates and a memory estimate.
  • Trace token-based admission and continuous batching.
  • Explain private prefix reuse and cancellation.
  • State retry/stream behavior after worker failure.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design an LLM inference platformWhat is the difference between prefill and decode?Recall first, then reveal

Prefill processes the prompt; decode generates subsequent tokens using retained attention state.

Read the prompt; write the continuation.

Return to lesson
Design an LLM inference platformDoes a prefix cache store the answer?Recall first, then reveal

No. It reuses compatible prompt computation; new output tokens still require generation.

Reuse the beginning, generate the ending.

Return to lesson
Design an LLM inference platformWhy is QPS insufficient?Recall first, then reveal

Requests differ in input tokens, output tokens, duration, and KV memory.

Count tokens and live context.

Return to lesson

Final revision

Summary and interview notes

An inference platform limits token work, waiting and memory before generation starts. Workers schedule prompt processing alongside ongoing token generation and own the KV blocks those computations use. Reuse only compatible authorized prefixes, wait for active computations before freeing memory, and bill from durable usage records.

Remember these points

  • Isolated prefill/decode throughput gives lower bounds, not proof of mixed-worker capacity.
  • The KV estimate depends on architecture, head count, element size and live sequence length in addition to model weights.
  • Prefix sharing requires compatible scoped identity, immutable shared blocks and safe reference lifetimes.
  • Cancellation stops future scheduling before in-flight work drains and memory becomes reusable.
  • Saving cumulative usage counters prevents repeated reports from duplicating charges. Finalizing the attempt prevents late reports from adding charges for work left unrecorded at the crash.

Interview tips

  • Before choosing hardware, calculate input/output tokens per second, active requests as arrival rate × mean service time (Little’s law), and KV bytes per retained token.
  • Explain who owns local capacity when two routers both see the last free slot.
  • Test cancellation during a kernel and a worker crash between a usage report and the next token.

Important qualifications

  • vLLM configuration changes over time; pin and load-test a runtime/model combination rather than relying on rolling-document defaults.
  • KV quantization and speculative decoding are workload-dependent optimizations with compatibility and quality requirements.
  • Output fencing prevents stale events being accepted; it does not by itself stop a GPU kernel or reclaim its memory.

Technical references

Practice marks stay in this browser.