Interview problem: provide an internal inference API for interactive assistants and offline jobs. Teams choose from approved model versions; the platform enforces tenant budgets, predictable latency, isolation and controlled upgrades.
All quantities below are interview assumptions. Benchmark the actual model, accelerator, runtime, precision, context lengths and output distribution before buying capacity.
1. Clarify scope
Ask whether external APIs are allowed, which models and modalities must be served, how traffic bursts, whether tenants require dedicated hardware, and what happens when demand exceeds the budget. Assume two approved text models, shared GPUs for ordinary tenants, and dedicated pools for contractual isolation. Training and arbitrary user-uploaded model code are out of scope.
Functional requirements
- Accept authenticated generation requests with a model alias, input, output limit and deadline.
- Support streaming responses, cancellation and an explicit terminal status.
- Enforce per-tenant request, concurrent-request, token and spending limits.
- Publish approved versions, canary changes and roll back routing.
- Run asynchronous batches separately from interactive work.
- Attribute usage to tenant, application and model revision without logging prompt text by default.
Non-functional requirements
- Target p95 time to first token below one second for the agreed short-prompt class at admitted load.
- Target p95 request-level time per output token below 50 ms for that class. Measure complete-request latency separately.
- Target 99.9% monthly availability for admitted interactive requests; report admission rejection separately so shedding does not hide poor service.
- Preserve tenant boundaries across caches, logs and adapters.
- Bound queue time and memory; reject work that cannot meet its deadline.
- Keep one failed replica from exhausting the surviving fleet.
2. Estimate tokens and memory
Assume a 100 requests/s peak, 2,000 input tokens and 300 output tokens per request.
| Quantity | Calculation | Design consequence |
|---|---|---|
| Prefill demand | 100 × 2,000 = 200,000 input tokens/s | Benchmark prefill separately |
| Decode demand | 100 × 300 = 30,000 output tokens/s | Single-stream speed is not fleet throughput |
| Mean in-flight work at 10 s | 100 × 10 = 1,000 requests | Queue and KV capacity matter |
| Raw 8B weights at two bytes | 16 GB, decimal | Excludes KV, activations, kernels and runtime |
| Example KV per stored token | 2 × 32 layers × 8 KV heads × 128 values × 2 bytes = 131,072 bytes | Architecture-specific, about 128 KiB/token |
| 100 sequences × 2,300 tokens | About 28.1 GiB of raw KV | Can exceed weight memory |
With a measured 2,500 output tokens/s per replica at the target latency, twelve replicas are the arithmetic throughput floor. At a chosen 70% utilization ceiling, ceil(30,000 / (2,500 × .7)) = 18 replicas. After losing one replica, check 17 × 2,500 × .7 = 29,750, which falls short: use at least nineteen under these assumptions. This still does not prove the prefill, KV-memory or multi-model constraints fit.
3. Start with a baseline
Read diagram source
flowchart LR
C[Application] --> A[Authenticated API]
A --> L[Bounded request queue]
L --> E[One model engine]
E --> C
A --> U[(Usage records)]
Use one model revision and a fixed output limit. Measure prefill and decode, queue wait, cancellation and useful completions before adding distributed scheduling.
4. Find the failures and justify repairs
| Baseline failure | Change | Benefit | Cost or limit |
|---|---|---|---|
| Long prompts block short interactions | Separate workload classes; test chunked prefill | More predictable interactivity | Scheduling and fairness tuning |
| One tenant consumes all KV memory | Token-aware admission and per-tenant concurrency | Contains noisy neighbors | Rejects some otherwise valid work |
| Idle batch jobs compete with live traffic | Independent batch pool or lower-priority preemptible work | Protects interactive SLOs | Lower utilization or restart cost |
| A rollout changes outputs unexpectedly | Versioned alias and shadow/canary comparison | Limits exposure | Extra inference and evaluation |
| Cancellations leave GPU work running | Propagate cancellation to the engine | Recovers capacity | Races with completion require cleanup |
Continuous batching improves scheduling opportunities. It does not remove memory limits or guarantee each request's latency.
5. Detailed architecture
Read diagram source
flowchart TD
C[Applications] --> G["Gateway<br/>identity and tenant policy"]
G --> B["Admission<br/>budget and token reservation"]
B --> Q["Scheduler<br/>deadline and tenant-fair queues"]
Q --> R["Router<br/>pinned model revision"]
CFG[(Approved configuration)] -.-> G
CFG -.-> R
R --> P1[Interactive GPUs]
R --> P2[Dedicated GPUs]
R --> P3[Batch GPUs]
P1 --> S["Streaming relay<br/>backpressure and cancellation"]
P2 --> S
P3 --> ART[(Batch results)]
S --> OUT[Client response]
P1 -.-> U[Usage and telemetry]
P2 -.-> U
P3 -.-> U
U --> SET[Reconcile reservations]
The request path pins an approved model revision. The following control loop publishes those approved revisions and adjusts capacity; it does not execute on every output token.
Read diagram source
flowchart LR
REG[(Signed artifacts)] --> LOAD[Load and readiness]
LOAD --> CAN[Canary and evaluation]
CAN --> CFG[(Approved configuration)]
OBS[Queue and token telemetry] --> CAP[Capacity controller]
CAP --> POOL[GPU pool capacity]
The data plane reads an approved configuration snapshot without depending on the control plane for every token. Security revocation needs a defined fast path; an indefinitely stale allow decision is not acceptable.
6. API and data contracts
POST /v1/generations accepts request_id, model_alias, input, max_output_tokens, deadline_ms and stream. Derive tenant identity from authentication. Return the pinned model_revision and a stream with ordered event IDs, a finish reason and final usage when available.
| Record | Key fields | Invariant |
|---|---|---|
| Deployment | model revision, tokenizer revision, runtime digest, precision, pool | An alias resolves to an approved compatible bundle |
| Admission | tenant, request ID, reservation, deadline, state | A request cannot reserve unlimited output |
| Usage | request ID, attempt ID, measured tokens, provider/engine result | Retries are attributable; final settlement is idempotent |
| Route | alias, eligible pools, rollout weights, policy version | A fallback must satisfy the same policy |
Idempotency can prevent duplicate request creation. It does not recreate an interrupted nondeterministic stream unless outputs are durably retained. State that boundary to the client.
7. Trace the request
- Authenticate, validate context/output limits, and select eligible model versions.
- Reserve a bounded token or monetary allowance before admission.
- Choose a pool using workload class, queue age and cache locality without violating tenant isolation.
- Pin the model/runtime configuration; the scheduler allocates KV blocks and admits compatible work.
- Relay output with backpressure. A slow or disconnected client triggers cancellation after a bounded grace period.
- Record measured usage, reconcile uncertain termination and release unused reservations.
Prefix reuse requires compatible tokens, positions, model/adapter state and isolation policy. See KV and prefix caches. A cache hit is a performance optimization, not proof of authorization.
8. Failure handling and operations
| Failure | Response | Evidence to monitor |
|---|---|---|
| GPU OOM | Fail bounded request; inspect lengths and admission; quarantine unstable replica | KV utilization, admitted tokens, OOM by revision |
| Replica dies before output | Retry only within the original deadline and attempt budget | Retry amplification and cold-start latency |
| Replica dies mid-stream | Emit interrupted status; offer explicit restart | Partial completions, lost tokens and user-visible errors |
| Control plane unavailable | Use last approved nonexpired policy where permitted | Configuration age and revocation lag |
| Region unavailable | Route only to authorized regions with tested spare capacity | Failover latency, quota and isolation |
Use vLLM's documented metrics as one implementation reference; metric names and semantics are runtime/version-specific. Monitor queue time, TTFT, per-request token latency, throughput and rejection by workload class. GPU utilization alone cannot explain service quality.
9. Cost and alternatives
For an illustrative nineteen replicas at $3/replica-hour, 19 × 730 × $3 = $41,610/month for that compute line. Add networking, storage, capacity for other models, observability, evaluation and people. Compare full cost per successful request with an approved hosted API; these are assumed prices, not vendor quotes.
| Choice | Prefer when | Tradeoff |
|---|---|---|
| Shared pool | Similar workloads and acceptable logical isolation | Better utilization, more fairness work |
| Dedicated pool | Contractual isolation or predictable sustained demand | Higher idle cost |
| Tensor parallelism | One model or latency target requires several GPUs | Communication cost and larger failure unit |
| More independent replicas | Model fits and aggregate traffic grows | Replicated weights; simpler request isolation |
| Prefill/decode separation | Measured interference justifies it | KV transfer and two capacity controllers |
10. Interview follow-ups
Q1: Can you size this service from requests per second alone?
Sample answer: No. Input/output lengths, active sequences, architecture-specific KV memory and latency targets determine the work. I benchmark representative token distributions and failure reserve, then check each independent bottleneck.
Follow-up: What if average output length doubles? Decode demand approximately doubles before changes in scheduling, memory pressure or queue behavior are considered.
Q2: Would you retry a disconnected stream transparently?
Sample answer: Only before any output has been delivered, and within a bounded contract. After partial output, a fresh generation may differ and duplicate content. I expose interruption or replay a retained stream using event IDs if the product requires resumability.
Follow-up: Does setting a random seed solve replay? Not as a portable service guarantee across concurrent execution, runtimes and model revisions.
Q3: What is the strongest closing decision?
Sample answer: Begin with token-aware admission, a measured interactive pool and a separate batch path. Add cache-aware routing and disaggregated serving only when traces show an economic or latency benefit. Demonstrate capacity after a replica failure and safe behavior when the client cancels.
Final summary and notes
| Remember | Interview evidence |
|---|---|
| Tokens before replicas | Prefill, decode and KV calculations |
| Admission before overload | Deadlines, quotas and reservation limits |
| Versions before rollout | Model, tokenizer, runtime and policy together |
| Outcomes before utilization | Useful completions and latency by class |
Tip: Draw cancellation and usage settlement. They distinguish an operated inference service from a GPU box on a diagram.