Atlas · skill

Classification

Classification predicts a discrete label or a distribution over labels from an observation. The skill includes defining classes, learning decision boundaries and selecting thresholds that match error costs. It also requires evaluating uncertainty and ambiguous cases, because a model's most likely class is not automatically an acceptable decision.

conceptSupervised Learning

What it is

A classifier maps features to class scores or probabilities and then, when needed, to a label. Binary, multiclass and multilabel tasks require different target representations and evaluation. The learning objective encourages correct distinctions from labeled examples, while regularization and representation shape the resulting boundary. Probabilities may need calibration before being used as risks. Thresholding turns continuous evidence into an action and can trade precision against recall. Classification competence therefore spans both the statistical model and the definition of the labels: inconsistent or overlapping categories limit what any fitted boundary can mean.

What the work involves

Create clear labeling instructions and check disagreement, missing labels and class prevalence. Choose features available at decision time and a validation design that separates relevant entities or periods. Compare a simple baseline, inspect the confusion matrix and tune thresholds on development data according to operational cost. Assess calibration and error slices, including an abstention or review path where uncertainty matters. The output should specify the class definitions, score interpretation and decision rule, enabling reviewers to distinguish a prediction from the action taken because of it.

Illustrative example

In an illustrative document-routing system, messages are assigned to billing, technical support or general inquiries. The team reviews ambiguous examples and introduces a manual-review path rather than forcing every message into a confident category. Evaluation examines which confusions delay customers, not only overall accuracy. A threshold for automatic routing is chosen on development data, and the remaining messages go to reviewers. The final test measures both routing quality and the fraction requiring review.

Limits and common mistakes

Classes can reflect annotation conventions rather than natural categories. Imbalanced data make accuracy misleading, and probability outputs are not necessarily calibrated. A closed-set classifier may confidently label an example outside all known classes. Changing prevalence or workflow can require threshold revision. Classification differs from clustering, which discovers groups without these target labels, and from regression on a continuous outcome. Check ambiguous, out-of-scope and costly-error cases before assuming a good average score supports fully automatic decisions.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10