Hypothetical interview scenario. Traffic, quality margins, GPU throughput, staffing and project costs are planning assumptions. Model documentation and quoted API prices have primary-source links.
Interview focus: Determine whether a specialized student model is worth building, then design the data, evaluation and serving controls needed to operate it.
1. Define distillation and the business decision
Knowledge distillation trains a student model using information produced by a teacher model. The teacher may provide probability distributions, target outputs or other permitted supervision. The student is often smaller, but size reduction is not part of the definition. In this design, a compact student learns specific customer tasks from checked teacher answers. It does not become a copy of every teacher capability.
The customer sends 8 million requests a month and currently spends $50,000 on a frontier-model deployment. About 90% of requests are recurring classification, extraction, summarization and three kinds of triage. Traffic grows by an assumed 18% per quarter. There is one ML engineer with part-time platform support. Customer data must remain in an approved region.
Start with the decision: Can simpler rules, a classifier, prompt changes, caching or a cheaper current hosted model meet the same requirements? A historical $50,000 bill is not evidence that distillation is the cheapest solution today.
Functional requirements
- Identify which tasks are eligible for a student, with a tested route for unsupported or difficult requests.
- Collect permitted customer examples, retain lineage and separate training, development and independent evaluation data.
- Generate teacher labels, check their correctness and record human-review coverage honestly.
- Train and version the student with its tokenizer, prompt format, inference configuration and data manifest.
- Evaluate the complete routed system against independent task outcomes and the best practical baseline.
- Run isolated shadow comparisons, then a bounded canary release with rollback.
- Monitor quality, task mix, full cost and latency; refresh or retire the student when evidence warrants it.
- Support customer data restrictions, deletion obligations and approved model-routing changes.
Non-functional requirements
- Quality: allow at most a 2% relative decline in an agreed higher-is-better task metric, with separate hard safety and customer-slice requirements. For 90% baseline accuracy, the floor is
90% × 0.98 = 88.2%; this differs from an 88% floor under a two-percentage-point margin. - Latency: clarify the requested p95 350 ms target. Propose it for short classification outputs; independently validate extraction and summary targets below. Do not quietly label first-token latency as completed-answer latency.
- Residency: training, inference, fallback, storage, backups, traces and reviewer access must all satisfy the customer's actual contract.
- Availability: preserve a tested fallback and enough warm capacity for a worker failure. If no permitted fallback exists, return an explicit unavailable/deferred result.
- Cost: compare full recurring cost and investment payback with a currently qualified alternative.
- Maintainability: one ML engineer cannot also supply thousands of domain-review hours or continuous on-call coverage. Name the required support and reviewer capacity.
Scope: one customer's authorized dataset and task routes. Sharing a base model across customers does not authorize pooling their training examples. Keep fresh policy facts in retrieval or tools instead of training them into weights on every update.
2. Begin with the least expensive viable baseline
| Option | Advantage | Limitation to test |
|---|---|---|
| Rules or a supervised classifier | Fast, cheap and inspectable for bounded labels | Insufficient for open-ended generation; still needs labels |
| Smaller hosted model | Little serving infrastructure to operate | Quality, regional eligibility, provider limits and recurring charges |
| Prompted local model | Tests regional control before training | GPU capacity and prompting may be sufficient already |
| Customer-trained student | Learns recurring behavior and output conventions | Training rights, label quality, refresh cost and operational responsibility |
| Frontier route for every request | Broad capability with simpler routing | May be unnecessarily expensive for some tasks |
Run these alternatives on the same task definitions. Compare cost per successful task, including retries and human corrections, rather than token price alone. A model that is cheaper per call but requires more failed retries may lose.
Contemporary candidates, not assumed winners
A local teacher candidate is Qwen3.8-27B; a compact student candidate is Qwen3.5-9B. Both publish Apache 2.0 artifacts. Their language backbones use hybrid attention, and both include vision components. Confirm the precise training/quantization support for the chosen engine; generic support for an architecture does not certify a customer-trained checkpoint. A text-only deployment may still load extra components depending on the implementation. Qwen3.8-27B, Qwen3.5-9B.
GPT-6 Luna is a current hosted comparison for focused tasks; its public model page lists $0.10/M input and $0.50/M output tokens at standard short-context rates. It is an inference baseline here, not a claim that Luna supports fine-tuning. GPT-6 Luna model documentation.
For a hosted teacher, verify the intended output-training rights before collecting labels. OpenAI's agreement restricts training competing models and defines specific permitted exceptions; Anthropic's commercial terms also restrict competing-model training without express approval. Owning an output does not cancel those restrictions. Do not assume a commercial student serving customer traffic qualifies for an exception intended for non-distributed classifiers. A rights-approved local teacher is the baseline for this scenario. OpenAI Services Agreement, Anthropic Commercial Terms.
3. Size traffic, latency and hardware
Eight million requests over 30 days average 8M / 2,592,000 ≈ 3.09 requests/s. At an assumed 20× burst, the service sees about 61.7 requests/s. With 90% routed directly to the student, peak student demand is approximately 55.6 requests/s. The difficult 10% may require substantially more tokens and time per request.
| Task | Proposed completed-response target | What makes it plausible or questionable |
|---|---|---|
| Classification | p95 below 350 ms for bounded inputs and at most eight output tokens | A classifier or measured low-latency serving path; queue and network time still count |
| Structured extraction | p95 below 2 seconds for the stated bounded input and about 100 output tokens | At an assumed 80 output tokens/s, decoding alone is 1.25 seconds |
| Summary | p95 below 6 seconds for about 300 output tokens | At 80 tokens/s, decoding alone is 3.75 seconds |
These are proposed requirements to agree with the interviewer, not benchmark results. Reasoning tokens, long input prefill and queueing can increase latency. If 350 ms truly applies to a 300-token completed summary, it would require over 857 output tokens/s before other delays; change output scope or architecture instead of asserting that an arbitrary 9B model meets it.
Suppose load testing eventually demonstrates 20 requests/s per replica at the required task mix and latency, and the plan uses 70% of that capacity: 14 requests/s per replica. Four surviving replicas support 56 requests/s. Provision five warm replicas to retain that capacity after one fails. A replica may occupy one or multiple GPUs; the cost example below explicitly assumes one. Validate availability-zone placement and correlated failures separately.
At 18% quarterly growth, the monthly workload would become 8M × 1.18^4 ≈ 15.51M after four quarters. Do not price a fixed five-replica fleet as sufficient forever.
4. Evolve from a pilot into a production pipeline
A pilot starts with a small rights-approved sample, teacher outputs and manual validation, then trains one model and compares it against the alternatives. Only add a durable pipeline after the pilot shows a useful gain.
| Pilot flaw | Production repair | Cost or tradeoff |
|---|---|---|
| Random row split | Group by source document/conversation and split by suitable time/customer boundaries | Smaller effective independent set, but less leakage |
| Teacher response treated as truth | Task validators, evidence checks and representative human audits | Reviewer effort and excluded/corrected examples |
| Only train loss tracked | Independent task/slice evaluation of the serving configuration | More evaluation time |
| All uncertain requests run through both models | Pre-route known unsuitable tasks; bound escalation and shadow samples | Routing evaluation and policy maintenance |
| Training launched on whatever GPU is available | Region-bound job admission, scoped storage and artifact lineage | Lower capacity flexibility |
| Model weights updated in place | Immutable release manifest and atomic route update | Storage and warm rollback capacity |
Read diagram source
flowchart TB
TRAFFIC[Approved customer requests] --> TRACE[Region-bound trace store]
TRACE --> SPLIT[Deduplicate and group<br/>training, development and assessment]
SPLIT -->|Training partition| LABEL[Approved teacher<br/>versioned evidence and label rules]
SPLIT --> DEV[Development cases for selection]
SPLIT --> HOLD[Restricted expert-labeled assessment]
LABEL --> CHECK[Automated validation<br/>expert sampling and corrections]
CHECK --> DATA[Accepted dataset manifest]
DATA --> TRAIN[Budgeted regional training jobs]
TRAIN --> REG[Signed model and tokenizer artifacts]
REG --> SELECT[Candidate selection]
DEV --> SELECT
SELECT --> EVAL[Independent task and slice evaluation<br/>deployed precision and routing]
HOLD --> EVAL
EVAL -->|Pass| SHADOW[Isolated sampled shadow]
SHADOW --> CANARY[Bounded customer-approved canary]
CANARY --> ROUTE[Versioned serving routes]
ROUTE --> STUDENT[Regional student pool]
ROUTE --> BACKUP[Tested permitted fallback]
STUDENT --> METRICS[Quality, latency and complete cost]
BACKUP --> METRICS
METRICS --> TRACE
Data and service contracts
| Record | Required fields | Why it matters |
|---|---|---|
| Trace | Tenant, task, source/version, consent/use policy, region, timestamps and outcome | A log entry is not automatically training-authorized |
| Dataset item | Input, target, source group, split, teacher/version, checks and review status | Distinguishes proposed, accepted, corrected and human-reviewed labels |
| Dataset manifest | Item digests, filters, rights version, splits, counts and exclusions | Training can be reproduced and affected artifacts traced |
| Training run | Base/tokenizer, dataset digest, recipe, seed, hardware, metrics and checkpoints | Avoids an unexplained weight file |
| Release | Weight/adapter digest, compatible base, tokenizer, engine, precision, template and evidence | Evaluation must match serving |
| Routing decision | Request ID, tenant, task, policy/version, primary route and escalation reason | Supports fair quality and cost attribution |
| Model invocation | Unique invocation ID, request ID, primary/fallback/shadow role, tokens and billed cost | Two actual calls are two charges, not duplicate records |
Use POST /training-runs with immutable dataset/recipe IDs and an idempotency key. The job scheduler independently checks rights, region and budget. Use POST /inference with a server-authenticated tenant and a known task type; the client cannot select another tenant's model by supplying its name. GET /releases/{id} reports approval and compatibility evidence, not merely training completion.
A practical stack is versioned regional object storage, a relational metadata catalog, PyTorch with a qualified FSDP or DeepSpeed setup, and vLLM for tested serving configurations. Pick the sharding strategy after measuring training memory. Stacking multiple sharding systems without a supported integration adds failure modes.
5. Curate labels without pretending the teacher is always right
- Check collection and training rights before sampling. Keep only the necessary fields.
- Group related traces and remove exact and near-duplicates across splits. A paraphrase of the same contract belongs with that contract.
- Sample task and language strata. Use real traffic to understand demand and synthetic examples to fill deliberate gaps; verify both.
- Attach the applicable source/policy snapshot. The same question can legitimately have a different answer after a policy change.
- Generate a proposed teacher answer. Validate schema, enumerations, source support, permitted actions and internal consistency.
- Human-review a representative sample and uncertain/high-risk items. Correct or exclude known errors and investigate bad batches.
- Publish a versioned dataset with measured error rates and an explicit record of what humans actually reviewed.
A 5% spot-check of 800,000 examples is 40,000 reviews. At 90 seconds each, it requires 1,000 reviewer-hours. It does not human-approve the other 760,000 examples. Neither teacher precision nor a small audit yields a fixed guaranteed student precision.
A concrete labeling example
Input: “The invoice total is correct, but the same invoice appears twice in the payment queue.” Expected task: classify an operations ticket. The target is duplicate_invoice, supported by the duplicate-queue evidence. A short permitted rationale can say “The issue is duplication, not an incorrect amount.”
Schema validation alone would accept incorrect_amount if it is another valid label. An evidence-aware validator or expert must establish that the label is correct. An ambiguous ticket should retain an unresolved label or route for clarification; forcing it into the most common category creates misleading training data.
This example function enforces admission metadata, not semantic correctness. A trusted pipeline supplies the checks after examining the actual content. accepted_under_batch_policy deliberately does not mean every item was human-reviewed.
def training_pair_status(pair, *, tenant, region):
if pair.get("tenant") != tenant or pair.get("region") != region:
return "reject_scope"
if pair.get("training_rights_valid") is not True:
return "reject_rights"
if pair.get("split") != "train" or pair.get("holdout_overlap") is not False:
return "reject_evaluation_overlap"
checks = pair.get("checks", {})
for name in ("schema", "evidence", "policy", "privacy"):
if checks.get(name) is not True:
return "hold_validation"
if pair.get("human_status") in ("incorrect", "uncertain"):
return "hold_expert_resolution"
if pair.get("human_status") == "correct":
return "accepted_human_reviewed"
if pair.get("human_status") == "not_reviewed" and pair.get("batch_approved") is True:
return "accepted_under_batch_policy"
return "hold_review_policy"
Recheck current rights before a job starts and before publishing its artifact. Immutable lineage records history; it cannot turn revoked permission into current permission. Keep access policy outside the model's control.
6. Choose the training objective and size the experiment
| Method | Supervision | Best use and limitation |
|---|---|---|
| Output/sequence distillation | Teacher-generated target text | Works without logits; inherits errors and sample selection bias |
| Distribution distillation | Teacher probabilities over aligned outputs | Supplies richer relative preferences; requires usable distribution access and compatible alignment |
| Rationale supervision | Permitted explicit explanations or solution steps | May help specific tasks; additional labels/tokens and quality checks |
| Ordinary supervised fine-tuning | Human or authoritative task labels | May remove the need for a teacher entirely |
Classical distillation matches softened teacher/student probabilities; temperature controls how concentrated those distributions are. Output-only instruction distillation is not the same objective. DistilBERT and TinyBERT are specific encoder studies, not evidence that every small generative model recovers 92–98% of a frontier model's quality. Original distillation paper, DistilBERT, TinyBERT.
The rationale experiments in Distilling Step-by-Step demonstrate task-specific benefits; they do not require exposing a provider's hidden internal reasoning. Use permitted explicit explanations or annotated steps. A student can learn from such supervision while the deployed classifier returns only a label. Measure whether the gain offsets collection, training and runtime cost.
Start with a small supervised or LoRA pilot and a held-out development comparison. Use full fine-tuning only if its extra capacity is useful. Preserve general/rare-task examples where relevant, and evaluate forgetting. Do not select epochs by training loss alone.
| Training quantity | Worked assumption |
|---|---|
| Accepted examples | 800,000 |
| Tokens/example | 1,000 input + 150 target = 1,150 |
| Tokens over three epochs | 800,000 × 1,150 × 3 = 2.76B |
| Pilot aggregate training throughput | 20,000 tokens/s across four GPUs |
| One run | 138,000 seconds = 38.33 wall-clock hours = 153.33 GPU-hours |
| Three experiments at $3/GPU-hour | $1,380 |
| Add 25% compute retry allowance | $1,725 |
The throughput and hourly rate are assumptions to replace with a measured pilot and regional quote. Input, target, padding, packing and loss-masking affect actual training work. At 800,000 pairs, a mandatory week on eight high-end GPUs is not a law of distillation. Conversely, a tiny compute estimate does not include the people needed to create trusted labels.
Quantize only after measuring the deployed result
For a nominal 9B language model, weight bytes alone are about 18 GB in BF16, 9 GB in INT8 or FP8, and 4.5 GB in idealized INT4. Scales, metadata, vision components, activations, allocator headroom and attention/recurrent state add memory. A fourfold reduction in weight bytes does not imply a fourfold throughput gain.
Benchmark the exact checkpoint, engine, GPU kernels, input/output distribution and load. Weight precision and KV/state precision are separate decisions. A format supported for one architecture or GPU may not be supported for another. Validate quality after quantization and include long-tail cases. vLLM quantization support.
7. Evaluate the product, then shadow and canary
An initial 1,800-case human-labeled assessment is a planning set, not a universal adequate sample. Use separate development cases for prompt/training selection and an independent final set. Group by document/conversation and check overlap with all training and teacher-labeling inputs. Human majority vote does not make an ambiguous answer true; adjudicate disagreement and record exclusions.
| Dimension | Suitable evidence | Common mistake |
|---|---|---|
| Classification | Per-class precision/recall and task cost, with explicit positive labels | Micro-average hides rare critical classes |
| Extraction | Field accuracy plus complete-record correctness | Valid JSON counted as correct |
| Summary/triage | Source-supported facts, omitted conditions and expert task rubric | Teacher agreement treated as ground truth |
| Routing | Full system outcomes at each fallback rate | Comparing teacher's hard cases with student's easy cases |
| Privacy and security | Cross-tenant, memorization and prohibited-action checks | Assuming redaction proves anonymity |
| Performance | End-to-end latency under realistic load, token counts and failures | Reporting only successful or short requests |
Use paired comparison and uncertainty for the agreed relative decline, with separate hard requirements. A two-point absolute margin from another chapter cannot be silently substituted. Reusing a judge across model versions requires calibration; consult evaluation gates.
In shadow mode, the student answers a permitted copy, but the user still receives the existing system's answer. Disable writes, notifications and payment tools in the copy. Both invocations consume resources. Randomly sample within a budget rather than duplicating all traffic by default.
In a canary release, a small real cohort receives the student route. An illustrative ramp is 5%, 20%, 50%, then up to 90%, advancing only after sufficient quality and operational evidence. Stratify or randomize assignment consistently; watch each important tenant/task/language and delayed outcomes. Thumbs-up response rate is selected feedback, not an unbiased quality estimate. Keep a representative paired audit sample separate from hard-case fallback traffic.
Read diagram source
sequenceDiagram
participant G as Authenticated task gateway
participant P as Routing policy
participant S as Approved student
participant T as Permitted fallback
participant V as Result validation
G->>P: Tenant, task, region and deadline
alt In qualified student scope
P->>S: Bounded request under release manifest
S->>V: Proposed answer and usage
alt Valid for delivery
V-->>G: Answer with recorded route
else Invalid or unsupported output
V->>P: Escalation reason and remaining budget
P->>T: Only if permitted and useful before deadline
T-->>V: Proposed fallback answer
V-->>G: Validated answer or explicit unresolved result
end
else Outside student scope
P->>T: Direct qualified fallback route
T-->>V: Proposed answer
V-->>G: Validated answer or explicit unresolved result
end
A model's self-reported confidence is not a calibrated routing probability. Pre-route known unsupported tasks, use measurable checks and calibrate any learned router on outcomes. A fallback started after the whole deadline is spent cannot rescue the original latency target. Do not execute downstream business actions before selecting and validating one answer.
8. Make the investment and payback concrete
Compare the current hosted alternative
At 1,000 input and 100 output tokens per request, GPT-6 Luna's standard rates yield 0.0001 + 0.00005 = $0.00015/request, or $1,200/month for 8M requests. With the documented 10% regional-processing premium where available, this becomes $1,320 before tools, retries and other charges. The region/endpoint must meet the actual requirement; availability in a broad geography is not permission to process in a particular cloud region. Quality is unproven until evaluated. GPT-6 Luna pricing and regional notes.
Initial student investment
| Item | Derivation or explicit allowance | Cost |
|---|---|---|
| Local teacher labels and validation compute | Pilot allowance $0.01 × 800,000 accepted pairs | $8,000 |
| Human 5% audit and correction | 40,000 × 90 seconds / 3,600 × $60/hour | $60,000 |
| Independent assessment labels | 1,800 × 3 reviewers × 3 minutes / 60 × $80/hour | $21,600 |
| Training experiments | Three assumed-throughput scenario runs plus retry allowance | $1,725 |
| Engineering and integration | 400 hours × $120/hour | $48,000 |
| Initial investment | Sum of listed work | $139,325 |
These are planning costs, not supplier quotes. Teacher validation is still fallible; count rejected and regenerated pairs in the pilot's effective cost per accepted pair. Three reviewers supply independent labels, and difficult adjudication beyond the allowance increases cost. Specialist annotation cannot be supplied by one ML engineer while also doing 400 integration hours.
Recurring student operating model
The historical average teacher cost is $50,000 / 8M = $0.00625/request. Assume the 10% fallback requests are three times as expensive as that average: $0.01875/request. Ten percent of requests therefore consumes $15,000, not $5,000.
| Monthly item | Assumption | Cost |
|---|---|---|
| Fallback teacher | 800,000 × $0.01875 | $15,000 |
| Student fleet | Five one-GPU replicas × 720 hours × $1.50/hour | $5,400 |
| Data, gateway and telemetry | Incremental allowance | $1,800 |
| Incremental operations | 8 hours/week × $120/hour × 52 / 12 | $4,160 |
| Refresh reserve | $24,000 every five months, including labeling/eval/deployment | $4,800 |
| Recurring included total | Serving, support and refresh reserve | $31,160/month |
The hourly rates differ between training and serving because this scenario budgets different GPU classes. Qualify the actual hardware and regional availability. The $24,000 refresh assumes a smaller incremental dataset; it cannot pay for a repeat of the entire $60,000 initial human audit plus other work.
Against the historical $50,000 model bill, treating common pre-existing platform costs as equal, normalized savings are $18,840/month, or $226,080/year. Initial payback is $139,325 / $18,840 ≈ 7.4 months after reaching the assumed steady state. A normalized twelve-month operating period minus investment leaves $86,755; actual first-year cash flow differs with the build/ramp delay and dates of refresh spending. Amortization is not a cash payment every month.
| Teacher request share | Recurring cost with the same student fleet | Savings versus $50,000 | Approximate payback |
|---|---|---|---|
| 5% | $23,660 | $26,340/month | 5.3 months |
| 10% | $31,160 | $18,840/month | 7.4 months |
| 20% | $46,160 | $3,840/month | 36.3 months |
At a fixed $16,160 non-teacher cost and $0.01875 per fallback request, the maximum teacher share for positive savings is about 22.56%. Higher traffic, a larger fleet, token mix and price changes alter that threshold. Requests that first run on the student and then escalate cost more than direct fallback; include their extra load. Shadow calls are additional too.
Decision under these assumptions: if the $1,200 hosted comparison passes the same quality and regional requirements, this student loses on cost. Distill only for a demonstrated advantage—such as required regional control or materially better task outcomes—or choose the simpler qualified alternative. There is no universal 200,000-request cutoff or guaranteed three-month payback.
9. Failure modes, repairs and tradeoffs
| Failure | Repair | Remaining cost or limitation |
|---|---|---|
| F1: Teacher improves | Compare current candidates on the same task set before relabeling | A new teacher is a reason to evaluate, not proof that retraining pays |
| F2: Distribution changes | Inspect new tasks/languages, policies and routing; temporarily use a qualified fallback | Input drift alone does not establish quality drift |
| F3: Teacher errors enter training | Find affected source groups, correct labels and retrain/evaluate the impacted release | A spot-check cannot remove every hidden error |
| F4: Fallback spend grows | Meter reason and tokens, reassess routes and budget | Do not force an unsafe student answer merely to hit a cost target |
| F5: Canary misses a rare failure | Add targeted tests and per-slice monitoring, revert affected routes | Small canaries provide limited evidence about rare events |
| F6: Residency violation | Block mismatched jobs/endpoints and inspect replicas, logs, backups and reviewers | Local weights alone do not ensure all processing stays local |
| F7: Rare-task forgetting | Include suitable replay examples, inspect per-task metrics and retain fallback | Oversampling needs weighting for population estimates |
| F8: Incorrect cost attribution | Reconcile unique provider invocations with primary/fallback/shadow tags | Both calls in a real comparison are legitimate costs |
Redaction reduces exposure; it does not prove anonymity or prevent memorization. Measure missed identifiers and context leakage, restrict access to artifacts, and retain a lineage-based deletion procedure for datasets, checkpoints, caches and backups. Deleting a row from the trace store does not remove its influence from an already trained model. The required retraining or other remediation depends on the documented data-use obligation.
10. Operational Considerations
| Signal | What to inspect | Action |
|---|---|---|
| Quality by task and customer slice | Independent outcomes, uncertainty and mislabeled examples | Revert a failing route, repair data or compare an alternative |
| Latency by output-length band | Queue, prefill, decode, retries and fallback duration | Reduce queueing/length or add measured capacity |
| Fallback share and token cost | Direct route versus post-student escalation | Review expensive patterns with a time-bound owner |
| Drift | Fixed encoder/bins, cohort, input length and sample size | Diagnose before choosing retraining |
| Refresh readiness | Rights, new labels, independent tests and available staff | Schedule only a bounded project with a useful expected gain |
| Successful task cost | Serving, audits, retries, refresh and human correction | Re-evaluate whether the student still earns its operating cost |
For example, Spanish policy questions increasing from 10% to 25% is a distribution change. First check whether their success actually fell and whether source policies, sampling or the measurement encoder changed. A new encoder can move a histogram without a user-facing regression. Repairing retrieval or routing may be faster than distillation.
Review comparable teacher/student evidence monthly and refresh candidates around a four-to-six-month planning cadence when warranted. A refresh proceeds through approved sample, corrected labels, training, independent assessment, shadow and canary; four calendar weeks is not a completion guarantee. Revert by switching a compatible model/tokenizer/template/router manifest, with enough fallback capacity already available.
Communicate material model-routing changes and obtain approval where required by the customer contract. Show what was evaluated and where the evidence is limited; never advertise “frontier quality” merely because a composite score is close.
Interview follow-ups
Q1. Is distillation just fine-tuning a smaller model?
It is training with teacher-provided supervision; fine-tuning is one way to carry it out. The teacher can provide probabilities or selected outputs. A student trained only on human labels is ordinary supervised training, even if it is smaller. I would specify the actual objective and what teacher information is available.
Q2. Why is a 5% human audit not enough to call the dataset clean?
It measures errors in a sample, with uncertainty and coverage limitations. The unreviewed examples may contain systematic errors. I would use automated checks on every item, targeted review of risky groups and representative auditing, then correct or exclude problematic batches. I would label review status accurately.
Q3. The student agrees with the teacher 98% of the time. Is that good enough?
Agreement can reproduce the teacher's mistakes. Compare with authoritative outcomes or independent expert labels, inspect severity and slices, and apply the declared margin. Also evaluate routing, quantization and latency. Agreement alone says nothing about training rights or production readiness.
Q4. Can a small generative model meet 350 ms for all these tasks?
Not by assumption. A short classification result may fit a carefully measured budget; a 300-token summary has a different decode lower bound. I would negotiate task-specific completion targets, consider a classifier for bounded labels and benchmark under realistic concurrency. First-token latency cannot substitute for completion latency.
Q5. What if fallback doubles from 10% to 20%?
Under the stated model, recurring cost rises from $31,160 to $46,160 and payback stretches to roughly 36 months. Inspect whether traffic changed, the student degraded or routing became overly conservative. Keep quality intact while considering a different model or stopping the project; a fallback cap cannot justify wrong answers.
Q6. Why shadow before canary?
Shadowing collects comparisons without changing the delivered answer, provided side effects are isolated. A canary measures the effect of actually serving the student. Neither replaces independent tests, and both need representative coverage. Count duplicate inference cost and watch delayed outcomes before expanding exposure.
Q7. Would you approve this project today?
I would first evaluate the cheaper hosted and local baselines. Under the worked numbers, the student is substantially more expensive than a qualified low-cost hosted model. I would approve a limited pilot only if regional constraints or measurable task quality create a plausible advantage, with a stopping criterion and a staffed review plan.
60-second interview answer
I would treat distillation as an investment decision before treating it as a training job. First compare current cheaper alternatives on the same tasks and regional requirements. If a student has an advantage, collect authorized examples, check teacher labels and preserve independent evaluation. Train a versioned model, test its deployed precision and routing, then shadow and canary with a qualified fallback. Count label review, warm GPUs, escalation and recurring refresh in the payback. Keep the option to retire the student when a simpler model meets the requirements more economically.
Remember: Compare alternatives → Authorize data → Check labels → Test the serving system → Roll out → Recalculate value.
Final notes: Relative percent and percentage points differ. A small weight file does not prove low latency. A sampled label audit is not universal approval. The best interview conclusion can be a well-supported decision not to distill.
Related: knowledge distillation, inference fundamentals, cost optimization, fine-tuning platform.