Principal Component Analysis
Principal component analysis projects data onto orthogonal directions ordered by captured variance. It supports compression, visualization and removal of redundant linear dimensions. The skill includes choosing scale, fitting the projection without leakage and assessing lost information, because variance preserved by an unsupervised projection is not necessarily the information a downstream task needs.
Also searchable as: PCA, principal component analysis (pca)
What it is
PCA centers a numeric representation and finds directions that successively maximize projected variance subject to orthogonality. It can be computed through singular-value decomposition, yielding component directions, scores and explained-variance quantities. Keeping fewer components produces a low-rank representation and a corresponding reconstruction. Feature scale determines which variables contribute most strongly; centering does not automatically standardize them. PCA does not use target labels and therefore optimizes a different objective from supervised feature selection. Component signs can be reversed without changing the represented subspace, and correlated or similarly strong directions can complicate interpretation of individual components.
What the work involves
Decide whether standardization matches the meaning of features and fit all preprocessing and the PCA projection on training data. Choose component count using reconstruction, variance and downstream validation rather than a universal cutoff. Inspect loadings and representative reconstructions, and measure behavior across relevant data slices. Save centering, scaling and components for consistent inference. The result should explain what variation the reduced representation retains and loses, including whether the compression improves a downstream model or simply makes the dataset easier to visualize.
Illustrative example
In an illustrative sensor project, several channels measure related physical variation. An analyst fits PCA on historical training readings and reconstructs held-out readings from a reduced set of components. A low-variance channel carries an important fault indicator, so retaining only the dominant components harms classification despite good average reconstruction. The analyst revises component selection and documents the distinction between compact representation and preserving information needed for the fault-detection task.
Limits and common mistakes
PCA captures linear variance, which can be dominated by scale, noise or nuisance variation. Low-variance directions can contain important target information. Components do not automatically correspond to interpretable or causal factors. Fitting on evaluation data leaks distributional information, and a two-dimensional plot can obscure structure lost in projection. PCA differs from nonlinear embedding methods and supervised selection. Evaluate reconstruction and task consequences, and avoid treating an explained-variance percentage as a complete measure of usefulness.
Prerequisites
- hardLinear Algebra
Eigenvectors, projections and covariance matrices are central to PCA.
Related skills
- → is subcategory of: Unsupervised Learning
Sources and further reading
- scikit-learn: PCA
Centering, solvers, components and explained variance.
- scikit-learn: Decomposition
Dimensionality reduction and matrix-factorization context.
Last updated: 2026-10-10