Classification
Classification predicts a discrete label or a distribution over labels from an observation. The skill includes defining classes, learning decision boundaries and selecting thresholds that match error costs. It also requires evaluating uncertainty and ambiguous cases, because a model's most likely class is not automatically an acceptable decision.
What it is
A classifier maps features to class scores or probabilities and then, when needed, to a label. Binary, multiclass and multilabel tasks require different target representations and evaluation. The learning objective encourages correct distinctions from labeled examples, while regularization and representation shape the resulting boundary. Probabilities may need calibration before being used as risks. Thresholding turns continuous evidence into an action and can trade precision against recall. Classification competence therefore spans both the statistical model and the definition of the labels: inconsistent or overlapping categories limit what any fitted boundary can mean.
What the work involves
Create clear labeling instructions and check disagreement, missing labels and class prevalence. Choose features available at decision time and a validation design that separates relevant entities or periods. Compare a simple baseline, inspect the confusion matrix and tune thresholds on development data according to operational cost. Assess calibration and error slices, including an abstention or review path where uncertainty matters. The output should specify the class definitions, score interpretation and decision rule, enabling reviewers to distinguish a prediction from the action taken because of it.
Illustrative example
In an illustrative document-routing system, messages are assigned to billing, technical support or general inquiries. The team reviews ambiguous examples and introduces a manual-review path rather than forcing every message into a confident category. Evaluation examines which confusions delay customers, not only overall accuracy. A threshold for automatic routing is chosen on development data, and the remaining messages go to reviewers. The final test measures both routing quality and the fraction requiring review.
Limits and common mistakes
Classes can reflect annotation conventions rather than natural categories. Imbalanced data make accuracy misleading, and probability outputs are not necessarily calibrated. A closed-set classifier may confidently label an example outside all known classes. Changing prevalence or workflow can require threshold revision. Classification differs from clustering, which discovers groups without these target labels, and from regression on a continuous outcome. Check ambiguous, out-of-scope and costly-error cases before assuming a good average score supports fully automatic decisions.
Prerequisites
Related skills
- → is subcategory of: Classical Machine Learning
- ← is subcategory of: Naive Bayes
- ← is subcategory of: Class Imbalance Handling
- ← is subcategory of: Logistic Regression
Sources and further reading
- scikit-learn: Model Evaluation
Classification metrics, confusion matrices and probability scoring.
- scikit-learn: Linear Model
Logistic and other linear classification methods.
Last updated: 2026-10-10