Atlas · skill

Mechanistic Interpretability

Mechanistic interpretability investigates how a neural network's internal components produce particular behaviors. It studies activations, learned features and computational circuits, using interventions to test hypotheses about mechanisms rather than relying only on input-output correlations or a model's verbal explanation of itself.

conceptExplainability & Fairness

What it is

A neural network distributes information across parameters and intermediate activations, often representing multiple concepts in overlapping directions. Mechanistic methods seek useful units of analysis such as attention heads, activation directions or features learned by sparse autoencoders. Researchers then examine which inputs activate these units and how modifying them changes outputs. Activation patching can replace an internal state from one run with a state from another to test a proposed causal contribution. Naming an interpretable feature is therefore a hypothesis about a representation; establishing a mechanism requires evidence that the relevant computation influences behavior under the tested conditions.

What the work involves

The practitioner starts with a narrowly defined behavior and a set of contrasting examples. They instrument model activations, identify candidate components and run interventions or ablations with suitable controls. A useful research artifact records the model version, input distribution, intervention location and measured output changes. Experiments should compare alternative explanations and check whether the proposed circuit generalizes beyond the discovery examples. Sparse features can help organize inspection, but reconstruction error and omitted computations must remain visible in the interpretation rather than disappearing behind convenient human labels.

Illustrative example

A researcher studies why a small language model completes a repeated name incorrectly. They compare correct and corrupted prompts, patch selected attention outputs between runs and measure changes in the next-token distribution. A candidate circuit is retained only when targeted interventions reproduce the predicted effect on new prompts. Visualizing attention alone would provide a clue, but would not establish that the highlighted connection causes the observed completion.

Limits and common mistakes

Current methods explain limited behaviors and model regions, not a complete account of a large model. Intervention choices can introduce artifacts, and apparently clean features may combine multiple meanings. A discovered mechanism can coexist with other paths that produce the same behavior. Mechanistic evidence should specify its scope and uncertainty; it is not, by itself, a certificate that a model is honest, aligned or safe in deployment.

Prerequisites

  • You inspect the weights and activations of a network.

  • Features and circuits are analyzed in activation space.

Sources and further reading

Last updated: 2026-10-10