K-Nearest Neighbors
K-nearest neighbors predicts from the labels or values of nearby stored examples under a chosen distance. It is a local, instance-based method rather than a compact fitted equation. The skill includes choosing a meaningful representation and neighborhood size and managing the quality and cost of retrieving neighbors at prediction time.
Also searchable as: KNN, k-NN, K Nearest Neighbors, KNN (K-Nearest Neighbors), k-nearest neighbors (knn)
What it is
For classification, neighbor labels contribute votes or weighted class evidence; for regression, neighbor outcomes are averaged or otherwise weighted. The number of neighbors controls locality, and distance weighting gives closer examples more influence. The model's behavior depends strongly on feature scale, the metric and the distribution of stored examples. Training largely establishes the reference data and any preprocessing, while prediction performs a neighbor query. Search implementations can be exact or use data structures suited to particular geometry. KNN is a predictive use of neighbors and should be distinguished from clustering, where the goal is to discover groups rather than infer known targets.
What the work involves
Scale or transform features inside the training pipeline and select a metric appropriate to their meaning. Tune neighborhood size and weighting on relevant validation data, checking rare groups and regions with sparse support. Estimate memory and query cost before choosing a search strategy. Inspect actual neighbors for representative errors so the notion of similarity can be challenged. The deliverable includes the reference data, preprocessing and prediction rule, with checks for behavior on observations far from the stored examples and a plan for updating the reference set.
Illustrative example
In an illustrative equipment classifier, an engineer predicts operating mode from sensor measurements using nearby labeled readings. Without scaling, one measurement's large numeric range dominates the distance. After correcting scale, the engineer compares small and larger neighborhoods and inspects ambiguous boundary cases. A reading far from all stored examples is sent for review rather than treated as well supported merely because the method can always identify the nearest available observations.
Limits and common mistakes
Distances can become less informative in high dimensions, and irrelevant features can overwhelm useful similarity. Small neighborhoods are sensitive to noise, while large neighborhoods blur local distinctions. Prediction cost and storage can grow with the dataset. Voting does not automatically yield calibrated probabilities or an out-of-distribution warning. Data duplicates and entity overlap can make validation optimistic. Check neighborhood quality, support and deployment latency rather than assuming that an intuitive distance guarantees an appropriate predictive relationship.
Prerequisites
- softLinear Algebra
Distance metrics and vector representations are easier to reason about with basic linear algebra.
Related skills
- → is subcategory of: Classical Machine Learning
Sources and further reading
- scikit-learn: Neighbors
Neighbor classification and regression, distance metrics and search methods.
Last updated: 2026-10-10