Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design a Multi-Tenant Model-Serving Platform

By Anup Rai8 min readReviewed September 2026

Interview problem: provide an internal inference API for interactive assistants and offline jobs. Teams choose from approved model versions; the platform enforces tenant budgets, predictable latency, isolation and controlled upgrades.

All quantities below are interview assumptions. Benchmark the actual model, accelerator, runtime, precision, context lengths and output distribution before buying capacity.

1. Clarify scope

Ask whether external APIs are allowed, which models and modalities must be served, how traffic bursts, whether tenants require dedicated hardware, and what happens when demand exceeds the budget. Assume two approved text models, shared GPUs for ordinary tenants, and dedicated pools for contractual isolation. Training and arbitrary user-uploaded model code are out of scope.

Functional requirements

  1. Accept authenticated generation requests with a model alias, input, output limit and deadline.
  2. Support streaming responses, cancellation and an explicit terminal status.
  3. Enforce per-tenant request, concurrent-request, token and spending limits.
  4. Publish approved versions, canary changes and roll back routing.
  5. Run asynchronous batches separately from interactive work.
  6. Attribute usage to tenant, application and model revision without logging prompt text by default.

Non-functional requirements

  1. Target p95 time to first token below one second for the agreed short-prompt class at admitted load.
  2. Target p95 request-level time per output token below 50 ms for that class. Measure complete-request latency separately.
  3. Target 99.9% monthly availability for admitted interactive requests; report admission rejection separately so shedding does not hide poor service.
  4. Preserve tenant boundaries across caches, logs and adapters.
  5. Bound queue time and memory; reject work that cannot meet its deadline.
  6. Keep one failed replica from exhausting the surviving fleet.

2. Estimate tokens and memory

Assume a 100 requests/s peak, 2,000 input tokens and 300 output tokens per request.

Quantity Calculation Design consequence
Prefill demand 100 × 2,000 = 200,000 input tokens/s Benchmark prefill separately
Decode demand 100 × 300 = 30,000 output tokens/s Single-stream speed is not fleet throughput
Mean in-flight work at 10 s 100 × 10 = 1,000 requests Queue and KV capacity matter
Raw 8B weights at two bytes 16 GB, decimal Excludes KV, activations, kernels and runtime
Example KV per stored token 2 × 32 layers × 8 KV heads × 128 values × 2 bytes = 131,072 bytes Architecture-specific, about 128 KiB/token
100 sequences × 2,300 tokens About 28.1 GiB of raw KV Can exceed weight memory

With a measured 2,500 output tokens/s per replica at the target latency, twelve replicas are the arithmetic throughput floor. At a chosen 70% utilization ceiling, ceil(30,000 / (2,500 × .7)) = 18 replicas. After losing one replica, check 17 × 2,500 × .7 = 29,750, which falls short: use at least nineteen under these assumptions. This still does not prove the prefill, KV-memory or multi-model constraints fit.

3. Start with a baseline

Architecture / visual model
flowchart LR C[Application] --> A[Authenticated API] A --> L[Bounded request queue] L --> E[One model engine] E --> C A --> U[(Usage records)]
Read diagram source
flowchart LR
 C[Application] --> A[Authenticated API]
 A --> L[Bounded request queue]
 L --> E[One model engine]
 E --> C
 A --> U[(Usage records)]

Use one model revision and a fixed output limit. Measure prefill and decode, queue wait, cancellation and useful completions before adding distributed scheduling.

4. Find the failures and justify repairs

Baseline failure Change Benefit Cost or limit
Long prompts block short interactions Separate workload classes; test chunked prefill More predictable interactivity Scheduling and fairness tuning
One tenant consumes all KV memory Token-aware admission and per-tenant concurrency Contains noisy neighbors Rejects some otherwise valid work
Idle batch jobs compete with live traffic Independent batch pool or lower-priority preemptible work Protects interactive SLOs Lower utilization or restart cost
A rollout changes outputs unexpectedly Versioned alias and shadow/canary comparison Limits exposure Extra inference and evaluation
Cancellations leave GPU work running Propagate cancellation to the engine Recovers capacity Races with completion require cleanup

Continuous batching improves scheduling opportunities. It does not remove memory limits or guarantee each request's latency.

5. Detailed architecture

Architecture / visual model
flowchart TD C[Applications] --> G["Gateway<br/>identity and tenant policy"] G --> B["Admission<br/>budget and token reservation"] B --> Q["Scheduler<br/>deadline and tenant-fair queues"] Q --> R["Router<br/>pinned model revision"] CFG[(Approved configuration)] -.-> G CFG -.-> R R --> P1[Interactive GPUs] R --> P2[Dedicated GPUs] R --> P3[Batch GPUs] P1 --> S["Streaming relay<br/>backpressure and cancellation"] P2 --> S P3 --> ART[(Batch results)] S --> OUT[Client response] P1 -.-> U[Usage and telemetry] P2 -.-> U P3 -.-> U U --> SET[Reconcile reservations]
Read diagram source
flowchart TD
 C[Applications] --> G["Gateway<br/>identity and tenant policy"]
 G --> B["Admission<br/>budget and token reservation"]
 B --> Q["Scheduler<br/>deadline and tenant-fair queues"]
 Q --> R["Router<br/>pinned model revision"]
 CFG[(Approved configuration)] -.-> G
 CFG -.-> R
 R --> P1[Interactive GPUs]
 R --> P2[Dedicated GPUs]
 R --> P3[Batch GPUs]
 P1 --> S["Streaming relay<br/>backpressure and cancellation"]
 P2 --> S
 P3 --> ART[(Batch results)]
 S --> OUT[Client response]
 P1 -.-> U[Usage and telemetry]
 P2 -.-> U
 P3 -.-> U
 U --> SET[Reconcile reservations]

The request path pins an approved model revision. The following control loop publishes those approved revisions and adjusts capacity; it does not execute on every output token.

Architecture / visual model
flowchart LR REG[(Signed artifacts)] --> LOAD[Load and readiness] LOAD --> CAN[Canary and evaluation] CAN --> CFG[(Approved configuration)] OBS[Queue and token telemetry] --> CAP[Capacity controller] CAP --> POOL[GPU pool capacity]
Read diagram source
flowchart LR
 REG[(Signed artifacts)] --> LOAD[Load and readiness]
 LOAD --> CAN[Canary and evaluation]
 CAN --> CFG[(Approved configuration)]
 OBS[Queue and token telemetry] --> CAP[Capacity controller]
 CAP --> POOL[GPU pool capacity]

The data plane reads an approved configuration snapshot without depending on the control plane for every token. Security revocation needs a defined fast path; an indefinitely stale allow decision is not acceptable.

6. API and data contracts

POST /v1/generations accepts request_id, model_alias, input, max_output_tokens, deadline_ms and stream. Derive tenant identity from authentication. Return the pinned model_revision and a stream with ordered event IDs, a finish reason and final usage when available.

Record Key fields Invariant
Deployment model revision, tokenizer revision, runtime digest, precision, pool An alias resolves to an approved compatible bundle
Admission tenant, request ID, reservation, deadline, state A request cannot reserve unlimited output
Usage request ID, attempt ID, measured tokens, provider/engine result Retries are attributable; final settlement is idempotent
Route alias, eligible pools, rollout weights, policy version A fallback must satisfy the same policy

Idempotency can prevent duplicate request creation. It does not recreate an interrupted nondeterministic stream unless outputs are durably retained. State that boundary to the client.

7. Trace the request

  1. Authenticate, validate context/output limits, and select eligible model versions.
  2. Reserve a bounded token or monetary allowance before admission.
  3. Choose a pool using workload class, queue age and cache locality without violating tenant isolation.
  4. Pin the model/runtime configuration; the scheduler allocates KV blocks and admits compatible work.
  5. Relay output with backpressure. A slow or disconnected client triggers cancellation after a bounded grace period.
  6. Record measured usage, reconcile uncertain termination and release unused reservations.

Prefix reuse requires compatible tokens, positions, model/adapter state and isolation policy. See KV and prefix caches. A cache hit is a performance optimization, not proof of authorization.

8. Failure handling and operations

Failure Response Evidence to monitor
GPU OOM Fail bounded request; inspect lengths and admission; quarantine unstable replica KV utilization, admitted tokens, OOM by revision
Replica dies before output Retry only within the original deadline and attempt budget Retry amplification and cold-start latency
Replica dies mid-stream Emit interrupted status; offer explicit restart Partial completions, lost tokens and user-visible errors
Control plane unavailable Use last approved nonexpired policy where permitted Configuration age and revocation lag
Region unavailable Route only to authorized regions with tested spare capacity Failover latency, quota and isolation

Use vLLM's documented metrics as one implementation reference; metric names and semantics are runtime/version-specific. Monitor queue time, TTFT, per-request token latency, throughput and rejection by workload class. GPU utilization alone cannot explain service quality.

9. Cost and alternatives

For an illustrative nineteen replicas at $3/replica-hour, 19 × 730 × $3 = $41,610/month for that compute line. Add networking, storage, capacity for other models, observability, evaluation and people. Compare full cost per successful request with an approved hosted API; these are assumed prices, not vendor quotes.

Choice Prefer when Tradeoff
Shared pool Similar workloads and acceptable logical isolation Better utilization, more fairness work
Dedicated pool Contractual isolation or predictable sustained demand Higher idle cost
Tensor parallelism One model or latency target requires several GPUs Communication cost and larger failure unit
More independent replicas Model fits and aggregate traffic grows Replicated weights; simpler request isolation
Prefill/decode separation Measured interference justifies it KV transfer and two capacity controllers

10. Interview follow-ups

Q1: Can you size this service from requests per second alone?

Sample answer: No. Input/output lengths, active sequences, architecture-specific KV memory and latency targets determine the work. I benchmark representative token distributions and failure reserve, then check each independent bottleneck.

Follow-up: What if average output length doubles? Decode demand approximately doubles before changes in scheduling, memory pressure or queue behavior are considered.

Q2: Would you retry a disconnected stream transparently?

Sample answer: Only before any output has been delivered, and within a bounded contract. After partial output, a fresh generation may differ and duplicate content. I expose interruption or replay a retained stream using event IDs if the product requires resumability.

Follow-up: Does setting a random seed solve replay? Not as a portable service guarantee across concurrent execution, runtimes and model revisions.

Q3: What is the strongest closing decision?

Sample answer: Begin with token-aware admission, a measured interactive pool and a separate batch path. Add cache-aware routing and disaggregated serving only when traces show an economic or latency benefit. Demonstrate capacity after a replica failure and safe behavior when the client cancels.

Final summary and notes

Remember Interview evidence
Tokens before replicas Prefill, decode and KV calculations
Admission before overload Deadlines, quotas and reservation limits
Versions before rollout Model, tokenizer, runtime and policy together
Outcomes before utilization Useful completions and latency by class

Tip: Draw cancellation and usage settlement. They distinguish an operated inference service from a GPU box on a diagram.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design an Image and Video Generation Platform
NEXT LESSONDesign a Multi-Tenant Vector Search Service →

Explore the diagram