Information Theory
Information theory measures uncertainty and the relationship between probability distributions. In machine learning it explains entropy, cross-entropy and divergence, and helps interpret compression and predictive losses. Competence means understanding what these quantities measure, choosing an appropriate representation and avoiding claims that a lower information-theoretic loss guarantees better task performance.
What it is
Entropy summarizes the uncertainty of a distribution, while cross-entropy measures the average coding or prediction cost when one distribution is represented by another. Kullback–Leibler divergence describes a directional mismatch between distributions; mutual information measures statistical dependence through shared information. These quantities connect probabilistic modeling, coding and learning objectives. A language model's token loss is a cross-entropy over its chosen tokenization, and perplexity is a transformed version of that loss. The units depend on the logarithm base. Information theory describes properties of distributions and representations, not the meaning, truthfulness or social value of the messages being modeled.
What the work involves
A practitioner should identify which distribution is treated as the reference, how probabilities are estimated and whether comparisons use the same support and preprocessing. Derive the relationship between a likelihood objective and its information-theoretic interpretation before using it as a model-selection metric. Account for smoothing when empirical probabilities are zero. For text models, document tokenizer and evaluation data so reported losses are comparable. The resulting analysis explains what uncertainty or mismatch was measured and why that quantity matters for the intended application.
Illustrative example
Imagine an illustrative autocomplete system evaluated on two collections of technical documents. The analyst computes average token cross-entropy on held-out text and checks that both runs use the same tokenizer and masking policy. A lower value suggests better prediction of those token sequences. The analyst then separately examines suggested completions for usefulness and factual correctness, because confidently predicting common wording does not establish that a completion answers the user's technical question.
Limits and common mistakes
KL divergence is generally asymmetric and is not a distance metric. Entropy estimates can be sensitive to sample size and representation, and mutual information can be difficult to estimate in high dimensions. Perplexities from different tokenizations should not be treated as directly interchangeable. Low uncertainty may reflect a narrow or repetitive dataset rather than broad competence. Cross-entropy evaluates probability assignments under a data distribution; it does not verify factual claims or capture every cost of a mistaken decision.
Prerequisites
Sources and further reading
- Dive into Deep Learning: Information Theory
Entropy, cross-entropy, KL divergence and mutual information in learning.
Last updated: 2026-10-10