Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Serving infrastructure: build a service around the model

By Anup Rai10 min readReviewed September 2026

Model serving is the system that accepts authorized inference requests, schedules model computation, returns results, and operates the workload within defined reliability and cost limits. An inference engine is one part of that system.

Define the interview contract

For a shared enterprise assistant, use the following illustrative requirements and confirm them with the interviewer.

Functional requirements

  1. Accept versioned chat requests for models each tenant is permitted to use.
  2. Stream responses, report completion or failure, and support cancellation.
  3. Attribute actual usage to the correct tenant and request.
  4. Roll out and roll back model/runtime versions without corrupting active streams.

Non-functional requirements

  1. Bound queueing and enforce per-tenant quotas under bursts.
  2. Meet agreed TTFT, inter-token and end-to-end latency targets for defined length ranges.
  3. Isolate tenant data, adapters, caches and logs according to policy.
  4. Survive a replica failure with explicit partial-response and retry behavior.
  5. Control cost per successful task while preserving quality.

Suppose expected traffic is 30 requests/second, with 2,000 input and 200 output tokens per request. The average demand is 60,000 input tokens/second and 6,000 output tokens/second. They are different work; do not combine them into one undifferentiated throughput number. If average time in the service is five seconds, Little's law suggests about 150 requests in the system at steady state. This includes waiting work, so it is not automatically the GPU active-batch size. Size burst headroom from a measured arrival distribution.

Start with one healthy serving pool

Architecture / visual model
flowchart TD C[Client] --> G[Gateway: identity, quota, deadline] G --> Q[Bounded admission queue] Q --> R[Router: eligible model and replica] R --> W[Runtime: prefill, decode, cache] W --> S[Streaming response and finish state] S --> C A[Versioned trusted model artifacts] --> W W --> M[Metrics, usage, redacted traces] G --> M D[Deployment controller: readiness and draining] --> W
Read diagram source
flowchart TD
  C[Client] --> G[Gateway: identity, quota, deadline]
  G --> Q[Bounded admission queue]
  Q --> R[Router: eligible model and replica]
  R --> W[Runtime: prefill, decode, cache]
  W --> S[Streaming response and finish state]
  S --> C
  A[Versioned trusted model artifacts] --> W
  W --> M[Metrics, usage, redacted traces]
  G --> M
  D[Deployment controller: readiness and draining] --> W

This baseline makes the trust and scheduling boundaries explicit. Add replicas when one measured pool cannot meet the workload, and add specialized pools only when isolation or resource differences justify them. A service registry must remove unhealthy replicas; readiness should exercise representative inference rather than only an open TCP port.

The Inference Gateway

The gateway accepts requests and applies policies before work reaches a model server.

Component Responsibility
Authentication and authorization Establish identity and enforce permitted models and actions.
Quotas and admission control Bound per-tenant usage and reject or defer excess work.
Model router Choose a model/version according to capability, rollout and availability policies.
Cache-aware routing Prefer a worker with reusable prefix state when useful; account for load and failure recovery.
Output handling Stream results and apply appropriate validation or risk controls. Filters are fallible.

Sticky routing is one cache-locality strategy, not a requirement for correctness. Prefix reuse must also respect isolation boundaries. Model-server scheduling and gateway routing are related but separate responsibilities.


Model Parallelism

If weights and request state do not fit or execute efficiently on one device, partition the work. A hypothetical 405-billion-parameter model needs 810 decimal GB for two-byte weight payload alone; this arithmetic does not identify a particular model or include cache and runtime overhead.

1. Tensor Parallelism (TP)

TP partitions operations within a layer across devices. It can reduce local matrix work, but collective communication and synchronization limit speedup. Fast interconnects help; NVLink is one option, not a logical requirement. Measure scaling on the actual topology. Four GPUs do not guarantee one quarter of single-GPU latency.

2. Pipeline Parallelism (PP)

PP puts groups of layers on different devices. A single sequence still traverses the stages in order; enough concurrent work or microbatches can keep stages busy. Imbalanced stages and empty pipeline slots create bubbles. PP can be useful within or across nodes, often combined with TP, depending on memory and network constraints.

3. Data and expert parallelism

Data parallel replicas handle different requests and increase fleet capacity when a model replica already fits. For MoE, expert parallelism distributes expert weights and routes token activations to the devices owning selected experts. Expert imbalance and all-to-all communication are additional concerns. “Only a few experts are active per token” does not mean all other expert weights disappear from storage.


Multi-GPU Orchestration

An orchestrator places replicas and manages lifecycle; the model runtime schedules tokens inside a replica. For example, KubeRay manages Ray workloads on Kubernetes. Gloo is a communication library in common ML usage, not an interchangeable Kubernetes operator.

Use heterogeneous devices only with an explicit placement and compatibility plan. Scaling signals can include token backlog, queue age, cache pressure, TTFT, output-token rate, and replica health. CPU utilization alone is often insufficient; KV utilization alone can also mislead when prefixes are reusable or the bottleneck is compute.

Cold start includes scheduling, image pull, weight transfer, initialization, optional compilation, and warmup. Measure each stage, cache trusted artifacts, and verify readiness with representative inference. Quantized versus unquantized artifacts and storage bandwidth affect loading; no universal 15–20-second startup follows from a particular image strategy. Keep warm capacity when startup exceeds the response deadline.


Streaming and Long-Lived Connections

SSE and WebSockets are common streaming transports; non-streaming and asynchronous batch interfaces also exist. Both L4 and L7 load balancers can support long-lived connections when configured appropriately. Check idle timeouts, buffering, drain behavior, connection limits, and disconnect propagation.

An ordinary L7 proxy does not inherently understand model end-of-sequence tokens. The application/runtime emits protocol completion and finish metadata. Route a new request or turn to an eligible replica; do not move a partially generated stream between unrelated workers without explicit state-transfer support. On deployment, drain existing requests and keep cancellation and billing attribution intact.


Compare inference engines on your workload

Compare supported capabilities, measured performance, and operating requirements separately. A documented feature may depend on a particular model, device, or configuration. Benchmark the configurations that meet your application's requirements.

vLLM: evaluate model coverage and serving behavior

vLLM documents continuous serving, prefix caching, LoRA, structured output, parallelism, and model-specific features. Check the selected release's supported-model and hardware matrices, then exercise the exact template, tool parser, dtype, and generation settings. Some disaggregation and cache-transfer combinations have compatibility limits; do not assume every feature composes with every other one. Start with vLLM features.

A useful benchmark has separate long-input, short-chat, structured-output, and burst slices. Record TTFT and inter-token latency alongside aggregate tokens/second, peak memory, failed requests, and output correctness. High throughput obtained by violating the application's latency target is not usable capacity.

SGLang: evaluate prefix reuse and scheduling

SGLang documents radix/prefix caching, structured generation, distributed serving, and hierarchical cache options. Shared-prefix workloads may benefit, but the value depends on prefix repetition, routing locality, memory pressure, and eviction. Compare cold and warm cache runs and use the same quality criteria as for other engines. The SGLang documentation and HiCache guide are capability references, not universal comparative speed claims.

A text-only benchmark cannot establish multimodal correctness or safety. Conversely, an unspecified historical vulnerability cannot justify declaring every current multimodal deployment unsafe. Assess the exact affected component and installed release using the project's advisory records and your exposure path.

TensorRT-LLM: evaluate the NVIDIA stack and backend

TensorRT-LLM provides NVIDIA-focused inference optimizations and multiple execution/deployment paths. The current quick start includes a PyTorch backend; it is incorrect to say every model always requires a multi-hour prebuilt TensorRT engine. Backend, model, hardware, precision, and feature support determine the setup and optimization work. See the TensorRT-LLM quick start.

Measure any conversion/build time, warmup, version compatibility, and operational work for the chosen path. A team with a stable model and NVIDIA expertise may justify additional tuning; a team changing architectures frequently may value faster iteration. These are evaluation criteria, not a binary rule choosing an engine for all organizations.

MoE-aware serving

MoE introduces expert-weight residency, token-to-expert routing, load imbalance, and inter-device communication. Active parameter count helps describe compute, while total parameters remain relevant to weight storage. Cache pressure and communication can prevent linear throughput scaling as batch size rises.

For a concrete experiment, hold model, precision, maximum context, and hardware fixed. Replay a representative mix at increasing offered load. Record throughput that still meets p95/p99 latency targets, expert-load distribution where observable, queueing, and communication time. Test a replica failure and a skewed workload. Do not assume that all engines schedule requests by shared expert activation or that a particular batch size always wins.

Decision framework: engine per measured workload

Architecture / visual model
flowchart TD A[Model, modality, hardware, and license constraints] --> B[Exclude unsupported configurations] B --> C[Pin runtime, kernels, model, template, and settings] C --> D[Replay equal-quality workloads at several loads] D --> E{Meets quality, latency, memory, and isolation gates} E -->|No| F[Repair configuration or reject candidate] E -->|Yes| G[Compare usable capacity and operating cost] G --> H[Canary, failure drill, and rollback]
Read diagram source
flowchart TD
    A[Model, modality, hardware, and license constraints] --> B[Exclude unsupported configurations]
    B --> C[Pin runtime, kernels, model, template, and settings]
    C --> D[Replay equal-quality workloads at several loads]
    D --> E{Meets quality, latency, memory, and isolation gates}
    E -->|No| F[Repair configuration or reject candidate]
    E -->|Yes| G[Compare usable capacity and operating cost]
    G --> H[Canary, failure drill, and rollback]
Workload What to compare Failure that the test should reveal
Mixed interactive chat Tail latency at sustained and burst load Queue growth hidden by good mean throughput
Shared long prefixes Cold/warm cache, locality, and eviction Claimed cache savings disappearing across replicas
Structured output/tools Schema adherence, parser behavior, grammar overhead Valid-looking responses with incorrect application semantics
Multimodal input Supported preprocessing and representative input sizes Unsupported formats, excessive resource use, or parser failures
Long-context generation Prefill/decode interference and KV memory Short requests stalling or admitted requests exhausting memory
MoE Expert balance, communication, and total weight residency Active-parameter arithmetic understating required hardware
Multiple adapters CPU/GPU cache limits and cold loads Incorrect adapter binding or cross-adapter prefix reuse

Operational posture

Pin the runtime, model revision, tokenizer/template, kernels, hardware class, precision, and serving policy in the release manifest. Check vLLM advisories and SGLang advisories for specific affected versions; inspect NVIDIA's release/security guidance for its stack. Record advisory ID, affected range, exposure, patched version, and validation rather than copying a generic minimum version forever.

A second engine is useful when it has a tested business purpose and the team can maintain it. Shadow or canary only after checking output contracts, data handling, and resource budgets. Keep a proven rollback; adding an untested engine during an outage can compound the failure.


Interview Questions

Q: Why is Tensor Parallelism preferred over Pipeline Parallelism for low-latency serving?

Strong answer: TP splits a layer across devices and can reduce its local compute time, but collectives, synchronization, and memory traffic limit scaling. PP splits layers across stages; one request traverses them in order while concurrent work can improve utilization. For low-latency serving I would compare both on the actual model and interconnect, potentially combining them. I would not promise latency divided by GPU count or assume PP is useful only across nodes.

Q: How do you handle "Noisy Neighbors" in a multi-tenant LLM cluster?

Strong answer: I enforce trusted per-tenant quotas and bounded queues at admission, then use supported fair scheduling or separate worker pools for required isolation. I account for token demand and KV pressure, not only request count. Runtime support differs, so I verify whether per-tenant scheduling exists rather than assuming the gateway controls individual GPU iterations. I test a flooding tenant during a provider failure and check that other tenants retain their service targets.


Find the limit, then change the design

Observed failure Change Benefit Cost or remaining flaw
One long prompt pauses active chats Chunked prefill or a separate long-input pool Protects ongoing output cadence New long requests may wait longer; separate pools use spare capacity less efficiently
Cache-local routing overloads one worker Balance reuse against queue and token load Avoids hot-worker tail latency Some prefix work must be recomputed elsewhere
Replica dies midstream Mark partial output and apply explicit retry rules Honest failure behavior A restart may regenerate different text; tools need independent idempotency
Scale-up arrives too late Warm reserve and earlier backlog signals Capacity exists before deadlines are missed Idle capacity costs money
Model fits only across devices Test TP, PP or combinations on actual links Meets memory and compute needs Communication, bubbles and failure domain increase
Prefill and decode need different resources Evaluate disaggregated serving Independent placement and scheduling KV transfer, routing, compatibility and recovery become additional systems

Disaggregated prefill/decode separates prompt processing and token generation into different serving resources. It can improve utilization or isolation for some workloads, but transferring and owning KV state adds network and recovery costs. Keep the integrated baseline until the measured benefit exceeds those costs; consult the chosen runtime's current compatibility matrix.

Additional practice

  1. A client retries after a partially streamed answer. What is safe to repeat? Generation can restart under an explicit user-visible contract; external actions require request IDs and independently enforced idempotency.
  2. Why not scale from GPU utilization alone? High utilization may be healthy or overloaded; combine queue age, token backlog, memory pressure, latency, failures and cold-start time.
  3. Does an OpenAI-compatible API make engines interchangeable? No. Verify templates, model behavior, tool parsing, streaming, error semantics and supported parameters.
  4. What is the benefit of separate prefill and decode pools? Independent resource allocation; quantify the KV-transfer and operating costs before adopting it.

Recall card and closing

Contract → admission → routing → runtime → recovery. State measured usable capacity, the first bottleneck, the smallest justified change, and the failure behavior. Preserve the model, runtime, template and policy revisions as one reviewable release.


Next: Cost Optimization Playbook

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← PagedAttention: allocate the cache as the sequence grows
NEXT LESSONAI cost optimization: improve the economics of a completed task →

Explore the diagram