K-Means Clustering
K-means partitions observations into a chosen number of groups by minimizing within-cluster squared Euclidean distances to centroids. The skill includes selecting features, scale and cluster count and assessing stability and interpretation. It is a useful geometric grouping method, but its objective does not establish that the resulting groups are meaningful categories.
Also searchable as: K-means, Kmeans, K Means
What it is
The algorithm alternates between assigning each observation to its nearest centroid and updating centroids from assigned observations. These steps reduce a within-cluster squared-distance objective until a stopping condition is reached. Initialization affects the solution because the objective can have multiple local minima. The cluster count is specified beforehand, and every observation receives an assignment in the standard formulation. K-means implicitly favors groups represented well by Euclidean centers. It differs from density-based methods that can discover irregular shapes and leave noise unassigned. Competence includes understanding what centroids and variance minimization mean for the selected feature representation.
What the work involves
Choose numeric features for which Euclidean distance is meaningful and scale them according to substantive importance. Compare cluster counts, initialization runs and stability under resampling. Inspect centroids, representative observations and borderline assignments, using internal criteria as diagnostic evidence rather than final proof of usefulness. Measure whether the groups support the intended analysis or downstream decision. The output should document preprocessing, cluster count and initialization, and explain how new observations are assigned and how changes in the input distribution will be detected.
Illustrative example
For an illustrative energy-use study, buildings are represented by normalized daily consumption profiles. An analyst fits several cluster counts and examines centroids for distinct usage patterns. One group contains unusually high overall consumption rather than a different profile, prompting reconsideration of normalization. After selecting a representation, the analyst reads building metadata and checks stability across months. The groups describe patterns under those choices; they do not prove inherent building types or causes of energy use.
Limits and common mistakes
K-means is sensitive to scale, outliers and initialization. Nonconvex shapes, strongly unequal densities and poorly chosen cluster counts can produce misleading partitions. Centroids may not correspond to an actual observation. Lower inertia follows from adding clusters and is not sufficient to choose their number. K-means is distinct from KNN prediction despite both using neighborhoods or distance. Report stability and ambiguity, and avoid equating a visually clean partition with a validated segmentation.
Prerequisites
- mediumCluster Analysis
Choosing K, scaling inputs and interpreting clusters depend on general clustering concepts.
Related skills
- → is subcategory of: Cluster Analysis
Sources and further reading
- scikit-learn: Clustering
K-means objective, initialization, cluster-number selection and geometric limitations.
Last updated: 2026-10-10