Atlas · skill

Knowledge Distillation

Knowledge distillation trains a student model to reproduce useful behavior or predictions from a teacher. Practitioners choose what information to transfer, which inputs expose it and how to measure the student's retained quality. The goal may be a smaller deployable model, a different architecture or a specialized policy.

conceptModel Compression

What it is

A teacher can provide probability distributions, intermediate representations or generated target responses. In classic classification distillation, softened class probabilities reveal relationships between alternatives that a single hard label omits. Temperature controls the softness of those distributions, and the student may combine a teacher-matching loss with ordinary labeled supervision. For language models, sequence-level distillation can train on teacher-generated answers, while token-level methods compare predictive distributions. These are different transfer mechanisms. The student learns through training on sampled inputs; it does not receive a literal compressed copy of every teacher parameter or necessarily inherit all of the teacher's capabilities.

What the work involves

Define the deployment constraint and choose a student capable of the required inputs and outputs. Build a transfer set that covers difficult cases, verify teacher targets and keep the final evaluation independent of both target generation and student tuning. Select the matching objective and any mixture with trusted labels. Evaluate the student against the teacher and a student trained without distillation, measuring task quality, latency and memory in the actual runtime. Preserve teacher provenance and generation settings. The result should identify the capabilities transferred successfully and the cases where the smaller model still needs escalation.

Illustrative example

An illustrative helpdesk classifier must run locally with limited memory. A larger teacher supplies class distributions for reviewed and unlabeled tickets. The student trains on these distributions plus trusted labels. Evaluation groups tickets by conversation and includes rare request types. The engineer finds that broad categories transfer well but subtle billing distinctions do not, and adds targeted labeled examples. Deployment testing checks the student's actual runtime rather than inferring speed from its parameter count.

Limits and common mistakes

Distillation transfers teacher mistakes and depends heavily on the input distribution. A student with insufficient capacity may imitate surface patterns while losing nuanced behavior. Synthetic outputs can introduce unsupported claims, and evaluating on teacher-generated references can favor imitation over correctness. Access to teacher probabilities, usage permissions and model interfaces can constrain the method. Distillation differs from quantization and weight merging because it learns a new model; compare quality and operating cost after that learning process, not merely model size.

Prerequisites

  • Distillation trains a student network to mimic a teacher network — both are neural networks requiring DL understanding

  • Evaluating distillation quality requires comparing student vs. teacher on meaningful metrics

Related skills

  • → is subcategory of: Model Fine-Tuning

Sources and further reading

Last updated: 2026-10-10