Inference is using a trained model to compute predictions or outputs. In an autoregressive language model, the service processes the prompt and then repeatedly predicts the next token from the available prefix. The useful performance question is where time and resources are spent across that complete request.
Follow one request
Read diagram source
flowchart LR
A[Client request] --> B[Validate and admit]
B --> C[Queue and tokenize]
C --> D[Prefill prompt]
D --> E[Sample next token]
E --> F[Stream token]
F --> G{Finished or cancelled?}
G -->|No| H[Decode with cached history]
H --> E
G -->|Yes| I[Release state and record usage]
Queueing, retrieval, tool calls, network buffering, and cancellation are application concerns as well as model-runtime concerns. A fast kernel does not fix an unbounded queue.
The Two Phases of Inference
Prefill processes input positions and constructs the attention state needed for generation. Known prompt tokens allow substantial parallel work. An engine may split a long prompt into chunks, so “the entire prompt is always one indivisible pass” is too strong.
Decode advances generated output using prior context, usually reusing stored keys and values. Conventional autoregressive decoding has a sequential dependency between newly generated tokens. Each step performs the relevant model-layer operations; it does not merely read one row of the weight matrix.
| Phase | What is known | Main work | Common pressure |
|---|---|---|---|
| Prefill | All input tokens | Projections, attention over prompt, cache creation | Compute, attention IO, prompt length, queueing |
| Decode | Input and already generated tokens | New-token projections, attention to retained history, sampling | Weight/cache bandwidth, compute at larger batches, scheduling |
For dense full attention, prompt attention computation includes a term quadratic in sequence length. FlashAttention changes how exact attention is computed and reduces memory traffic; it does not turn every full-attention operation into linear computation. Linear projections and feed-forward layers have different scaling. Review attention before quoting one complexity for the whole model.
Compute-bound and memory-bound are measured conditions
A workload is compute-bound when arithmetic capacity limits performance; it is memory-bandwidth-bound when movement of required data limits it. Prefill often has higher arithmetic intensity, while low-batch decode often has lower intensity. Batch size, architecture, context length, quantization, kernels, and hardware can change the limiting resource.
- Measure the actual request distribution and offered load.
- Separate admission/queue delay, prompt processing, and output generation.
- Inspect device utilization, memory bandwidth, cache pressure, and communication.
- Change the component associated with the observed limit.
- Repeat the same workload and compare both latency and correctness.
Do not optimize all workloads as if decode were always bandwidth-bound or prefill always compute-bound. Efficiently Scaling Transformer Inference provides a primary systems treatment.
Performance Metrics
| Metric | Definition to state | Measurement caution |
|---|---|---|
| TTFT | Time from the chosen request start to the first output token | Client and server clocks include different work |
| Inter-token latency | Time between successive emitted tokens | Buffering can hide or create visible bursts |
| TPOT | Average time per output token after the first, under a stated convention | Not necessarily equal to p95 inter-token latency |
| End-to-end latency | Time until the complete response is available | Includes output length and any other workflow steps |
| Throughput | Completed requests or tokens per time interval | State input versus output tokens and success criteria |
| Goodput | Work completed while meeting the required service targets | Definition and rejected/failed requests must be explicit |
Worked example: first token arrives after 0.8 seconds; 101 output tokens finish at 4.8 seconds. There are 100 intervals after the first token, so average TPOT is (4.8 − 0.8) / 100 = 0.04 seconds, or 40 ms. Shortening TTFT alone cannot remove the remaining four seconds of generation.
There is no universal requirement that every product needs TTFT below 200 ms or a complete response within two seconds. Agree targets for the task, then inspect p95/p99 as well as averages.
Match the intervention to the cause
| Observed problem | Candidate change | Cost or regression |
|---|---|---|
| Long repeated prompt dominates TTFT | Compatible prefix caching | Memory, eviction, eligibility, and isolation |
| Long new prompt disrupts existing streams | Chunked prefill | New request's TTFT and scheduling overhead |
| Weight traffic dominates decode | Supported quantization | Approximation and kernel compatibility |
| Low utilization with queued requests | Appropriate batching | Larger batches can worsen individual latency |
| Low-batch decode has spare compute | Speculative decoding | Drafting/verification overhead and acceptance rate |
| Fleet queue grows during bursts | Admission limits and measured capacity scaling | Rejections or delay; cold-start headroom |
FP8 is a family of numerical formats, not a guarantee of twice the speed with a fixed accuracy loss. Verify supported hardware and kernels, scaling, selected tensors, and task quality. The same caution applies to choosing an accelerator from peak FLOPS alone.
Interview practice
- Why can generation take longer than classification? It often requires many sequential output steps, while a classification head may produce its result in one forward computation. A generative classifier can still use decoding.
- Does prefix caching skip all prefill? It can reuse a compatible cached prefix; uncached suffixes and other request work remain.
- Why can throughput improve while users see worse latency? More batching or queueing can improve aggregate utilization while increasing wait or per-request time.
- Is decode always memory-bound? No. Batch, model, context, precision, and implementation determine the limit.
- Why report output length alongside latency? Generating 500 tokens is different work from generating 50; comparisons need equivalent tasks and output requirements.
- What happens on cancellation? Propagate it to the runtime, release request state, and record actual usage so abandoned work does not consume capacity indefinitely.
Recall card and closing
Queue → prefill → decode → deliver. Locate the delay, define the measurement boundary, and choose one intervention whose benefit can be tested under realistic load.