A KV cache stores attention keys and values for already processed tokens so autoregressive generation can reuse them. A prefix cache reuses compatible prompt-prefix computation across requests. An answer cache returns a previous result. These save different work and have different correctness requirements.
Calculate what is retained
For a conventional decoder with uniform layer shapes:
KV bytes ≈ 2 × layers × KV_heads × head_dimension
× retained_tokens × concurrent_sequences × bytes_per_element
The two represents keys and values. For 80 layers, eight KV heads, width 128, 128,000 tokens, and two-byte elements, one sequence uses 41.94 decimal GB, or 39.06 GiB, of ideal cache payload. Add weights, allocation metadata, workspace, and other model state. This is an illustrative architecture, not a named model specification.
The cache is a resource consumed throughout a request. A short prompt followed by a very long output can create pressure later, even if admission looked inexpensive. Account for maximum retained history and concurrency, and decide how to reject, preempt, or shorten work when capacity is exhausted.
Grouped-query attention changes the shape
Multi-head attention (MHA) uses multiple query, key, and value heads. Grouped-query attention (GQA) shares each KV head across a group of query heads. Multi-query attention (MQA) uses one KV head. These are model-architecture choices, not arbitrary serving toggles for any trained checkpoint.
| Example architecture | Query / KV heads | Relative ideal cache |
|---|---|---|
| MHA | 32 / 32 | 1 |
| GQA | 32 / 8 | 1/4 |
| MQA | 32 / 1 | 1/32 |
The ratios assume all other dimensions match. Quality depends on the trained model and task; heads are not assigned universal human roles such as “logic” or “creativity.” See the GQA paper.
Trace a prefix-cache lookup
Read diagram source
flowchart TD
A[Authorized request and model revision] --> B[Build compatible cache identity]
B --> C{Matching prefix available?}
C -->|Yes| D[Reuse its KV state]
C -->|No| E[Compute and optionally retain prefix]
D --> F[Process uncached suffix]
E --> F
F --> G[Decode new output]
Cache compatibility can depend on token IDs, model weights, position handling, adapter, modality processing, and runtime rules. A byte-identical human-readable document is not enough if its tokenization or model context differs. Namespace caches for the isolation requirements; an entry being present never grants access to its content.
Example: three requests have a 2,000-token shared policy prefix and different 100-token questions. A warm prefix cache may avoid repeating the 2,000-token prefill. It does not reuse the different questions or automatically return an old answer. Changing an early system instruction can invalidate reuse for later tokens.
Memory tiers and eviction
Where supported, a cache hierarchy can use GPU memory, CPU host memory, and backing storage. HBM is GPU memory, not a separate tier after “VRAM.” A host or disk hit must transfer data; compare that cost with recomputation. Include finite capacity, eviction, version invalidation, and worker failure. SGLang HiCache documents one concrete hierarchy.
Provider prompt caching is a billing contract too
API providers may implement automatic or explicit caching with model-specific minimum lengths, lifetimes, matching rules, and read/write/storage charges. Track reported cached-input usage instead of assuming all repeated text is discounted. Check the dated pricing reference before making a financial estimate.
With illustrative costs of $1 for an uncached prefix, $1.25 to create its cache entry, and $0.10 for each later hit, n uses cost:
uncached = n
cached = 1.25 + 0.10 × (n − 1)
Caching wins above about 1.277 uses, so two uses suffice under these assumptions. Real storage charges, expirations, misses, and incompatible releases change the threshold. Put stable material first only when doing so preserves the prompt's intended semantics.
Four ways to reduce cache or context cost
| Intervention | Resource it changes | Risk or cost |
|---|---|---|
| Paging and prefix sharing | Allocation waste and repeated compatible state | Metadata, eviction, identity and reference management |
| Cache quantization | Bytes per stored value | Quality change, scale metadata, kernel support |
| Architecture-level compression such as MLA | Representation of attention state | Requires compatible trained architecture and runtime |
| Retrieval or summarization | Information included in the prompt | Missing evidence, lost qualifications, summary errors |
None inherently extends the model's validated context window by a guaranteed multiplier. See the MLA explanation for the architecture distinction.
Compare long context with retrieval
For a stable 50,000-token document, cached long context is a reasonable baseline if the model can use it reliably and the user is allowed to see all of it. Retrieval may win when the corpus grows, permissions differ by passage, or selective evidence improves quality. Whole-document inclusion does not guarantee that the model notices every relevant clause.
Evaluate answer correctness, citation support, update frequency, cache hit rate, TTFT, output behavior, and complete cost. Combine retrieval with caching when a stable instruction prefix and changing evidence make that useful.
Interview practice
- What does the KV cache avoid? Recomputing earlier keys and values during conventional autoregressive generation; new tokens still require work.
- Why use KV-head count in the formula? GQA and MQA share keys and values across query heads, changing storage.
- Is a prefix hit an answer hit? No. It reuses prompt processing, while the new output is still generated.
- What invalidates reuse? Incompatible tokens, weights, adapters, positions, runtime configuration, or authorization boundaries.
- Why might a disk cache lose to recomputation? Transfer and lookup latency can exceed the saved compute for the workload.
- When is cached context better than retrieval? Only when the measured task, access, freshness, latency, and cost comparison supports it.
Recall card and closing
Shape → compatibility → locality → lifetime → economics. Explain the saved work and what remains, then show how the cache behaves after a model release, an access change, and a miss.