Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Quantization: budget memory without guessing quality

By Anup Rai5 min readReviewed September 2026

Quantization represents numerical values using a restricted set of levels, commonly to reduce the bits used for model weights, activations, or cached attention tensors. It introduces approximation. Whether that approximation improves usable serving capacity depends on the model, workload, hardware, and implementation.

The interview question is not simply “can we use four-bit?” It is “which tensors dominate the budget, which formats have efficient kernels, and what quality change is acceptable?”

Separate payload from total memory

For an illustrative model with eight billion stored weights:

Weight precision Ideal weight payload What is excluded
16 bits 16 decimal GB Scales, other tensors, buffers, cache, runtime
8 bits 8 decimal GB Same exclusions
4 bits 4 decimal GB Same exclusions; metadata can matter more
2 bits 2 decimal GB Same exclusions; quality may become harder to preserve

Use number of values × bits / 8 for the ideal payload. Then add quantization metadata, unquantized tensors, temporary workspace, allocator overhead, and the KV cache. A four-bit checkpoint fitting on disk does not prove the service fits in GPU memory at the required concurrency.

Weight-only quantization reduces weight storage. Weight-and-activation quantization also changes activation representation or computation. The storage dtype and arithmetic dtype are separate choices; some kernels dequantize low-bit weights before computation.

Choose a method for the workload

Method Mechanism Interview caution
NF4 A nonuniform 16-value codebook designed around normally distributed weights A distribution assumption, not a guarantee for every tensor
GPTQ Uses calibration and approximate second-order information to reduce layer-output reconstruction error Calibration coverage and kernel compatibility affect results
AWQ Uses activation information and scaling to reduce errors in important weight channels It does not simply retain a special 1% subset in high precision
FP8 Floating-point formats with eight bits and different exponent/fraction allocations Format, scaling, tensor selection, and hardware support must agree

The primary descriptions are QLoRA/NF4, GPTQ, and AWQ. These are different approaches, not interchangeable names for the same four-bit artifact.

A calibration set should represent expected inputs, lengths, and domains. A conversion that looks good on short English chat may regress on long technical documents. Compare the exact converted artifact against the higher-precision baseline using the same tokenizer, template, decoding configuration, and evaluation.

Follow a deployment decision

Architecture / visual model
flowchart TD A[Measure weights, cache, and workspace] --> B{Which budget fails?} B -->|Weight storage| C[Evaluate supported weight quantization] B -->|Long-context cache| D[Evaluate KV quantization and admission limits] B -->|Compute or scheduling| E[Profile kernels and batching first] C --> F[Quality and realistic load test] D --> F E --> F F --> G{Meets all limits?} G -->|Yes| H[Version artifact and canary] G -->|No| I[Revise precision, model, or capacity]
Read diagram source
flowchart TD
  A[Measure weights, cache, and workspace] --> B{Which budget fails?}
  B -->|Weight storage| C[Evaluate supported weight quantization]
  B -->|Long-context cache| D[Evaluate KV quantization and admission limits]
  B -->|Compute or scheduling| E[Profile kernels and batching first]
  C --> F[Quality and realistic load test]
  D --> F
  E --> F
  F --> G{Meets all limits?}
  G -->|Yes| H[Version artifact and canary]
  G -->|No| I[Revise precision, model, or capacity]

Memory reduction may improve concurrency by allowing more active requests. It may also add dequantization overhead without helping a compute-bound workload. Measure time to first token, time between tokens, throughput, tail latency, and cost per successful request separately.

File format and runtime are separate decisions

GGUF is a model-file format carrying tensors and metadata, used by llama.cpp and other tooling. It supports multiple quantization types; execution can involve CPU, supported GPU backends, and offloading. EXL2 is associated with ExLlamaV2 and supports allocating different bit widths toward an average bits-per-weight target.

A target of 4.5 bits per weight is an average storage objective, not a literal uniform 4.5-bit value or a total VRAM guarantee. Include scales and nonweight state. As checked in September 2026, the ExLlamaV2 repository is archived and directs development to ExLlamaV3. Evaluate maintained runtime support when choosing a new deployment rather than repeating old format rankings.

Pin the conversion tool and runtime versions. Validate supported architecture, tokenizer, tensor layout, device backend, and export fidelity. “GGUF” alone does not specify a speed, and “GPU format” alone does not establish the fastest end-to-end service.

KV cache: calculate the actual shape

For a conventional decoder cache, an ideal payload estimate is:

bytes = 2 × layers × KV_heads × head_dimension
        × cached_tokens_per_sequence × sequences × bytes_per_value

The factor two accounts for keys and values. Use KV heads, which may differ from query heads with grouped-query attention. Different architectures or cache layouts need their own calculation.

With 32 layers, eight KV heads, width 128, one sequence, and 8,192 tokens, BF16 gives 2 × 32 × 8 × 128 × 8192 × 2 = 1,073,741,824 bytes, or 1 GiB. Ideal eight-bit and four-bit payloads are 0.5 and 0.25 GiB, before metadata and unquantized components.

At two million tokens, the same shape would require about 262.14 decimal GB in BF16. This arithmetic does not imply the model supports such a context. Cache quantization may reduce memory but can affect quality, especially on long contexts. Concurrency still competes for compute and bandwidth even when cache capacity improves.

Post-training quantization versus QAT

Post-training quantization (PTQ) converts a trained model without a full quantization-aware training stage; some methods use calibration. Quantization-aware training (QAT) trains with simulated or explicit quantization effects so parameters can adapt to the target representation.

Try a supported PTQ baseline first when it can meet the requirement. QAT or distillation may recover useful quality when PTQ falls short, at the cost of data, training, evaluation, and a more involved artifact pipeline. No parameter-count threshold makes QAT universally mandatory.

Interview practice

  1. Does four-bit imply four times faster than sixteen-bit? No. Storage, memory bandwidth, compute kernels, and conversion overhead determine performance.
  2. Why can an eight-billion-weight model exceed 4 GB at four-bit? Metadata, unquantized tensors, cache, and runtime memory are outside the ideal weight payload.
  3. What is the difference between AWQ and GPTQ? They use different calibration-based methods to preserve useful behavior under low-bit weights; compare the deployed implementations.
  4. Is FP8 equivalent to FP16 with fewer fraction bits? Different FP8 variants have their own exponent/range and precision behavior, requiring appropriate scaling and support.
  5. When will weight quantization fail to solve OOM? When long-context cache, activations, buffers, or concurrency dominate the remaining memory.
  6. What should a release comparison include? Task quality, important slices, full memory, realistic tail latency, throughput, cold start, and artifact reproducibility.

Recall card and closing

Tensor → representation → kernel → workload → quality. Close with a measured memory breakdown and a quality budget. Avoid a speed claim based only on bits per weight.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Synthetic data: create examples for a specific learning gap
NEXT LESSONRLVR and GRPO: train against a checkable outcome →

Explore the diagram