Text Classification
Text classification assigns a predefined category or set of categories to a text unit. Practitioners define labels, select representations and evaluate prediction errors in the application context. The competence includes class imbalance, ambiguous examples and unsupported categories, rather than assuming every text can be forced into one useful label.
What it is
The input may be a message, sentence, document or selected passage, and its scope affects the meaning of the prediction. Single-label classification chooses one mutually exclusive class; multilabel classification can assign several. Models can use sparse lexical features with a linear classifier, contextual encoder representations or other learned mappings. Training targets express a labeling policy rather than a universal property of the text. Decision thresholds convert scores into output labels, especially in multilabel or rejection settings. Classification differs from sequence tagging, which labels spans or tokens, and from clustering, which discovers groupings without a fixed supervised inventory. The representation and annotation scheme jointly define the task.
What the work involves
Write label definitions and resolve overlaps before collecting training examples. Audit class frequencies, annotation agreement and source-specific artifacts. Establish a sparse-feature baseline and compare more complex models only against the same split and criteria. Hold out conversations, document families or time periods that could leak repeated content. Measure per-class precision and recall, confusion and relevant error costs; tune thresholds on development data. Inspect truncation and uncertain examples. The deliverable is a calibrated decision policy where needed, a versioned preprocessing and label mapping and evidence that the model generalizes to the texts that will actually be classified.
Illustrative example
An illustrative archive classifier separates technical reports, invoices and correspondence. A baseline performs well because file headers expose the categories, but deployment includes scans with missing headers. The developer evaluates body-only examples and groups template families across splits. They compare lexical and encoder models, inspect invoice-correspondence confusion and add an uncertain route for mixed documents. Final assessment checks the downstream review workload as well as category accuracy.
Limits and common mistakes
Aggregate accuracy can hide weak minority classes, while inconsistent labels limit what a model can learn. Models may rely on headers, author names or other shortcuts instead of relevant content. Long-document truncation and domain change can alter decisions. A high score does not automatically mean a calibrated probability, and fixed categories require explicit unknown handling. Sentiment and intent detection are particular classification tasks with additional semantics. Evaluate label policy, representation and decision threshold together.
Prerequisites
Related skills
- → is subcategory of: NLP
Sources and further reading
- Hugging Face Transformers: Text classification
Supervised sequence-label prediction, label mappings and inference.
- scikit-learn: Classification of text documents using sparse features
Lexical-feature baselines and inspection of document-classification shortcuts.
Last updated: 2026-10-10