Linear Regression
Linear regression estimates a continuous response through a linear combination of specified predictors. The skill includes selecting features and transformations, fitting coefficients and assessing assumptions, uncertainty and predictive error. Its simplicity supports inspection, but a linear coefficient is a conditional association unless a separate design justifies a causal interpretation.
Also searchable as: linear-regression
What it is
The ordinary least-squares formulation chooses coefficients to minimize summed squared residuals between observed and predicted outcomes. Linearity concerns the coefficients: transformed features and interaction terms can represent nonlinear patterns in original variables. Rank and correlated predictors affect whether coefficients are uniquely and stably estimated. Statistical uncertainty calculations need assumptions about errors and sampling, which are separate from the algebraic fitting criterion. Regularized variants change the objective to constrain coefficients and can improve prediction in suitable settings. Linear regression is therefore both a predictive model and a statistical framework, but its adequacy depends on the specification and intended interpretation.
What the work involves
Define the continuous outcome and inspect scale, functional relationships and data provenance. Fit a baseline with meaningful predictors, check residual patterns, influential observations and collinearity, and choose uncertainty calculations that respect dependence or unequal error variance. Validate predictions on relevant held-out cases. Report coefficient coding and transformations clearly and assess sensitivity to alternative specifications. The deliverable should separate parameter interpretation from predictive quality, enabling a reader to see which assumptions support an interval and which evidence supports future-use performance.
Illustrative example
For an illustrative building-energy model, an analyst predicts daily consumption from temperature and occupancy. Residual plots reveal curvature, so the analyst adds a justified temperature transformation rather than assuming the straight-line relationship is adequate. They validate on later days and inspect unusually influential holidays. The coefficient on occupancy is interpreted conditional on the included terms; a plan to reduce occupancy would require additional causal reasoning rather than simply reading the coefficient as a guaranteed saving.
Limits and common mistakes
Outliers, correlated predictors and misspecified functional form can undermine fitting or interpretation. Constant error variance is not guaranteed, and repeated observations can invalidate ordinary standard errors. A good average error may hide poor behavior in tails or unseen ranges. Regularization changes estimates and should not be described as ordinary least squares. Linear regression differs from logistic regression's probability link and discrete target. Check residuals and deployment relevance while avoiding causal claims based solely on fitted coefficients.
Prerequisites
- mediumStatistical Inference
Coefficient uncertainty, diagnostics and inference require statistical foundations.
Related skills
- → is subcategory of: Regression Analysis
Sources and further reading
- scikit-learn: Linear Model
Ordinary least squares, regularization and feature transformations.
- NumPy: Linear algebra
Least-squares solvers, rank and numerical stability.
Last updated: 2026-10-10