Atlas · skill

Logistic Regression

Logistic regression predicts class probabilities from a linear combination of features passed through a logistic or related multiclass link. The skill includes representation, regularization, coefficient interpretation and threshold choice. It is a transparent classification baseline, provided its probabilities and assumptions are checked for the population and decision in which it will be used.

Also searchable as: logistic-regression, LogisticRegression

conceptSupervised Learning

What it is

In binary logistic regression, a linear predictor determines log-odds and the logistic function maps it to a probability between zero and one. Fitting commonly minimizes a log-loss objective, optionally with regularization. Multiclass formulations can use a softmax model or other strategies depending on the implementation. The decision boundary is linear in the supplied features, though interactions or transformations can make it nonlinear in the original measurements. Coefficients describe conditional changes on the log-odds scale. Despite its name, logistic regression is generally used for classification, and a coefficient should not be treated as a causal effect without a separate identification argument.

What the work involves

Define the target class and encode features consistently, including scaling where regularization makes scale important. Choose a penalty and strength through development evaluation, inspect correlated predictors and confirm how categories are represented. Assess discrimination and calibration separately, then set an operational threshold using costs or capacity. Report coefficient interpretations only with their coding and model context. The deliverable includes a reproducible preprocessing and scoring pipeline plus an explanation of the probability and decision rule, rather than a list of coefficients detached from their assumptions.

Illustrative example

Imagine an illustrative registration-risk model using session length and prior visits. An analyst fits regularized logistic regression and checks calibration on later sessions. A nonlinear transformation of session length can represent a relationship that the raw linear term misses. The team selects a review threshold against available capacity. A positive coefficient for prior visits describes the fitted conditional association; it does not establish that encouraging extra visits would increase the modeled outcome.

Limits and common mistakes

Complete separation, strong collinearity and sparse categories can make coefficient estimates unstable. Regularization changes interpretation and needs suitable scaling. Linear log-odds may be an inadequate functional form, and probabilities may miscalibrate after prevalence shifts. Logistic regression differs from linear regression's ordinary continuous-response formulation and from a neural model with learned representations. Inspect convergence, residual or calibration behavior and influential cases, and distinguish a simple interpretable model from an automatically correct explanation.

Prerequisites

  • Probabilities, odds and likelihood are needed to understand training and calibrated outputs.

Related skills

Sources and further reading

Last updated: 2026-10-10