Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Knowledge distillation: transfer useful behavior, measure the loss

By Anup Rai5 min readReviewed September 2026

Knowledge distillation trains a student model using supervision produced by a teacher model. That supervision can be predictions, probability distributions, generated responses, or intermediate representations. The teacher is often larger or more capable, but “teacher” and “student” describe roles rather than required parameter counts.

In a Learnastra interview, frame distillation as a deployment decision: can a cheaper or more controllable model meet a specific quality target after learning from a stronger system? Do not infer a commercial model's undisclosed training recipe from its product name.

Define the transfer contract

Consider a document classifier whose large-model baseline is accurate but expensive at high volume. The student needs the correct category and an appropriate abstention on unfamiliar documents. It does not need to reproduce every capability or stylistic habit of the teacher.

  1. Define the task distribution, target quality, latency, and serving budget.
  2. Measure teacher errors and disagreement with reviewed labels.
  3. Choose the supervision the teacher can actually expose.
  4. Generate and filter training examples with provenance and version records.
  5. Train the student and evaluate it independently of teacher agreement.
  6. Compare total development and refresh cost with serving savings.
Architecture / visual model
flowchart LR A[Representative prompts] --> B[Teacher supervision] B --> C[Verification and filtering] C --> D[Student training] D --> E[Held-out product evaluation] E --> F[Serve student with monitored fallback]
Read diagram source
flowchart LR
  A[Representative prompts] --> B[Teacher supervision]
  B --> C[Verification and filtering]
  C --> D[Student training]
  D --> E[Held-out product evaluation]
  E --> F[Serve student with monitored fallback]

How Distillation Works

Supervision What the student learns from Access needed Main limitation
Hard labels or generated text A selected answer or token sequence Teacher output Hides uncertainty and alternative predictions
Soft distributions Relative probability across output classes or tokens Compatible probabilities or logits Output spaces must align; full distributions may be unavailable
Intermediate features Selected hidden representations Teacher activations and a mapping between layers/shapes More coupling and training complexity

For soft-label distillation, first turn logits into probabilities. With temperature T > 0:

p_teacher = softmax(z_teacher / T)
p_student = softmax(z_student / T)
L_soft = T² × Σ p_teacher × log(p_teacher / p_student)

This is a common scaled KL-divergence objective. It is not KL divergence applied directly to raw logits. A training recipe may combine it with supervised cross-entropy on reviewed targets. Temperature softens or sharpens the distributions; select it through evaluation rather than treating a fixed interval as universal. Hinton and colleagues introduced the influential soft-target formulation.

Worked classification example: a teacher assigns (invoice 0.70, receipt 0.25, contract 0.05). A hard target retains only “invoice.” A soft target also communicates that “receipt” is a closer alternative. If the document is actually a receipt, reproducing the teacher distribution still reproduces its mistake. Independent labels remain valuable.

For language models with different tokenizers, token-by-token distribution matching is not automatically well-defined. Text-response distillation is easier to apply across vocabularies; more specialized cross-tokenizer methods require an explicit alignment design.

Features, trajectories, and student mistakes

Feature matching requires access to the chosen activations, not merely a statement that the teacher is open-weight. Layers and dimensions may differ, requiring projection or correspondence choices. Matching hidden states does not prove that the student has inherited a faithful “conceptual map.”

Off-policy response distillation commonly trains on teacher-generated responses. On-policy distillation uses states or trajectories sampled from the current student and teacher feedback on those states. This can target the student's own mistakes, but requires additional teacher queries and a careful objective. The Thinking Machines on-policy distillation report is a useful concrete implementation; its speedups should not be generalized to every workload.

Filtered self-training and verifiable tasks

A model can generate candidate solutions, filter them with a checker, and train on accepted examples. This is a form of self-training or self-distillation. Avoid treating “self-distillation from proof” as the universally documented recipe for proprietary reasoning models.

  1. Generate diverse candidates for prompts with independently checkable answers.
  2. Run code in an isolated environment or use an appropriate answer/proof checker.
  3. Reject invalid, duplicated, contaminated, or unsupported examples.
  4. Train on accepted examples while retaining evaluation of the original task distribution.
  5. Check intermediate explanations separately if their correctness matters.

Passing a few tests is not a formal proof; a correct final answer can accompany flawed reasoning. DeepSeek-R1 documents reasoning-model training and distillation results, but those results do not establish another provider's implementation.

Quantization-Aware Distillation

Distillation can be combined with low-precision training or adaptation so a student learns to reduce errors introduced by quantization. A higher-precision teacher can provide a useful target, but quality recovery is an empirical outcome. Ordinary post-training quantization may already meet the requirement without another training loop. Compare quantization alone against quantization plus adaptation.

Cost, failure, and repair

Assume a student saves $0.004 per request and data preparation, training, and evaluation cost $24,000. Simple break-even is six million requests. Add retraining, teacher generation, fallback traffic, engineering, and hosting utilization before deciding. The numbers are exercise assumptions, not current API prices.

Failure Why it happens Repair and cost
Student copies confident errors Teacher outputs are treated as truth Reviewed labels and independent verifiers add data cost
Rare tasks regress Training overrepresents easy, common examples Stratified coverage and regression sets require curation
Student imitates long explanations Style is rewarded more than task success Short correct targets and task metrics may reduce verbosity
Savings disappear Fallbacks, long outputs, or low utilization dominate Measure full request cost and simplify routing
Dataset cannot be used as intended Rights or provider terms were not checked Confirm current permissions before generation and training

The final row is a procurement check, not a blanket claim about all providers' licenses. Record the permissions for the actual model, outputs, and training purpose.

Interview practice

  1. Must the student be smaller? No, though compression is a common motivation. The roles refer to the direction of supervision.
  2. Why use soft labels? They expose relative alternatives, provided the distribution is available, compatible, and useful.
  3. Does teacher agreement measure correctness? It measures imitation. Both models can agree on a wrong answer.
  4. What changes with on-policy distillation? Training visits the current student's states rather than relying only on teacher-generated trajectories.
  5. Can verified final answers teach flawed reasoning? Yes. Outcome verification does not establish every intermediate claim.
  6. When should the team skip distillation? When a simpler smaller-model baseline meets the target or expected savings do not repay development and maintenance costs.

Recall card and closing

Task → teacher signal → filter → student evaluation → lifecycle cost. State what is transferred, what may be lost, and how that loss will be detected. Continue with synthetic data for the dataset pipeline.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Preference learning: RLHF and DPO
NEXT LESSONSynthetic data: create examples for a specific learning gap →

Explore the diagram