Atlas · skill

Named Entity Recognition

Named entity recognition identifies text spans that refer to defined categories such as people, organizations or domain-specific items. The skill is designing an annotation scheme, locating exact boundaries and evaluating span and type errors. Recognizing a mention is separate from linking it to a database record or resolving every reference.

conceptText Understanding

What it is

NER maps tokens or character spans to entity labels under a chosen ontology. Systems may use sequence tagging, span classification, transition-based prediction, rules or combinations of these. Boundary conventions determine whether titles, modifiers and punctuation belong to an entity. Some tasks allow nested or overlapping mentions, whereas particular model implementations may not. Context helps distinguish a person's name from an ordinary word or an organization from a location. A recognized span remains a mention in the text; entity linking assigns a canonical identity, and coreference resolution connects mentions that refer to the same thing. These are neighboring operations with separate training and evaluation requirements.

What the work involves

Write labeling rules with positive examples and difficult boundary cases. Review annotation disagreements and audit whether the selected model supports overlapping spans. Keep whole documents or entity-rich source families together across splits. Train or configure the recognizer, preserve character offsets and compare exact span-plus-type precision and recall by category. Inspect missed entities, partial spans and ordinary phrases falsely labeled as names. If the output feeds redaction or a database, evaluate that downstream effect as well. The deliverable is a recognizer and annotation contract whose labels and coordinates can be used consistently on unseen text.

Illustrative example

An illustrative maintenance-note extractor recognizes equipment names and component identifiers. Annotators disagree about whether a preceding vendor name belongs in the equipment span, so the team resolves that rule before training. Evaluation holds out entire manuals and reports exact boundaries. A model identifies the device correctly but includes a neighboring date in some spans; reviewing highlights exposes this defect. Entity linking to the asset register is then tested as a separate stage.

Limits and common mistakes

NER depends on the label inventory and annotation conventions, so results from different schemes are not directly comparable. Unseen abbreviations and domain changes can reduce recall. Exact matching penalizes partial boundaries that may still be useful, while token accuracy can conceal missed rare entities. A detected organization name does not prove which organization it denotes. Distinguish recognition, linking and coreference, and verify source-text offsets before using predicted spans for redaction or structured records.

Prerequisites

No prerequisites.

Related skills

  • → is subcategory of: Natural Language Processing

Sources and further reading

Last updated: 2026-10-10