Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Inference: follow the request before optimizing it

By Anup Rai4 min readReviewed September 2026

Inference is using a trained model to compute predictions or outputs. In an autoregressive language model, the service processes the prompt and then repeatedly predicts the next token from the available prefix. The useful performance question is where time and resources are spent across that complete request.

Follow one request

Architecture / visual model
flowchart LR A[Client request] --> B[Validate and admit] B --> C[Queue and tokenize] C --> D[Prefill prompt] D --> E[Sample next token] E --> F[Stream token] F --> G{Finished or cancelled?} G -->|No| H[Decode with cached history] H --> E G -->|Yes| I[Release state and record usage]
Read diagram source
flowchart LR
  A[Client request] --> B[Validate and admit]
  B --> C[Queue and tokenize]
  C --> D[Prefill prompt]
  D --> E[Sample next token]
  E --> F[Stream token]
  F --> G{Finished or cancelled?}
  G -->|No| H[Decode with cached history]
  H --> E
  G -->|Yes| I[Release state and record usage]

Queueing, retrieval, tool calls, network buffering, and cancellation are application concerns as well as model-runtime concerns. A fast kernel does not fix an unbounded queue.

The Two Phases of Inference

Prefill processes input positions and constructs the attention state needed for generation. Known prompt tokens allow substantial parallel work. An engine may split a long prompt into chunks, so “the entire prompt is always one indivisible pass” is too strong.

Decode advances generated output using prior context, usually reusing stored keys and values. Conventional autoregressive decoding has a sequential dependency between newly generated tokens. Each step performs the relevant model-layer operations; it does not merely read one row of the weight matrix.

Phase What is known Main work Common pressure
Prefill All input tokens Projections, attention over prompt, cache creation Compute, attention IO, prompt length, queueing
Decode Input and already generated tokens New-token projections, attention to retained history, sampling Weight/cache bandwidth, compute at larger batches, scheduling

For dense full attention, prompt attention computation includes a term quadratic in sequence length. FlashAttention changes how exact attention is computed and reduces memory traffic; it does not turn every full-attention operation into linear computation. Linear projections and feed-forward layers have different scaling. Review attention before quoting one complexity for the whole model.

Compute-bound and memory-bound are measured conditions

A workload is compute-bound when arithmetic capacity limits performance; it is memory-bandwidth-bound when movement of required data limits it. Prefill often has higher arithmetic intensity, while low-batch decode often has lower intensity. Batch size, architecture, context length, quantization, kernels, and hardware can change the limiting resource.

  1. Measure the actual request distribution and offered load.
  2. Separate admission/queue delay, prompt processing, and output generation.
  3. Inspect device utilization, memory bandwidth, cache pressure, and communication.
  4. Change the component associated with the observed limit.
  5. Repeat the same workload and compare both latency and correctness.

Do not optimize all workloads as if decode were always bandwidth-bound or prefill always compute-bound. Efficiently Scaling Transformer Inference provides a primary systems treatment.

Performance Metrics

Metric Definition to state Measurement caution
TTFT Time from the chosen request start to the first output token Client and server clocks include different work
Inter-token latency Time between successive emitted tokens Buffering can hide or create visible bursts
TPOT Average time per output token after the first, under a stated convention Not necessarily equal to p95 inter-token latency
End-to-end latency Time until the complete response is available Includes output length and any other workflow steps
Throughput Completed requests or tokens per time interval State input versus output tokens and success criteria
Goodput Work completed while meeting the required service targets Definition and rejected/failed requests must be explicit

Worked example: first token arrives after 0.8 seconds; 101 output tokens finish at 4.8 seconds. There are 100 intervals after the first token, so average TPOT is (4.8 − 0.8) / 100 = 0.04 seconds, or 40 ms. Shortening TTFT alone cannot remove the remaining four seconds of generation.

There is no universal requirement that every product needs TTFT below 200 ms or a complete response within two seconds. Agree targets for the task, then inspect p95/p99 as well as averages.

Match the intervention to the cause

Observed problem Candidate change Cost or regression
Long repeated prompt dominates TTFT Compatible prefix caching Memory, eviction, eligibility, and isolation
Long new prompt disrupts existing streams Chunked prefill New request's TTFT and scheduling overhead
Weight traffic dominates decode Supported quantization Approximation and kernel compatibility
Low utilization with queued requests Appropriate batching Larger batches can worsen individual latency
Low-batch decode has spare compute Speculative decoding Drafting/verification overhead and acceptance rate
Fleet queue grows during bursts Admission limits and measured capacity scaling Rejections or delay; cold-start headroom

FP8 is a family of numerical formats, not a guarantee of twice the speed with a fixed accuracy loss. Verify supported hardware and kernels, scaling, selected tensors, and task quality. The same caution applies to choosing an accelerator from peak FLOPS alone.

Interview practice

  1. Why can generation take longer than classification? It often requires many sequential output steps, while a classification head may produce its result in one forward computation. A generative classifier can still use decoding.
  2. Does prefix caching skip all prefill? It can reuse a compatible cached prefix; uncached suffixes and other request work remain.
  3. Why can throughput improve while users see worse latency? More batching or queueing can improve aggregate utilization while increasing wait or per-request time.
  4. Is decode always memory-bound? No. Batch, model, context, precision, and implementation determine the limit.
  5. Why report output length alongside latency? Generating 500 tokens is different work from generating 50; comparisons need equivalent tasks and output requirements.
  6. What happens on cancellation? Propagate it to the runtime, release request state, and record actual usage so abandoned work does not consume capacity indefinitely.

Recall card and closing

Queue → prefill → decode → deliver. Locate the delay, define the measurement boundary, and choose one intervention whose benefit can be tested under realistic load.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← RLVR and GRPO: train against a checkable outcome
NEXT LESSONKV and prefix caches: reuse computation with clear boundaries →

Explore the diagram