Atlas · skill

spaCy

spaCy is a library for building language-processing pipelines with tokenization, trained linguistic components and rules. The competence is selecting compatible language assets, configuring dependencies and inspecting structured annotations. It supports efficient application workflows, but each component and rule still needs evaluation on the language and domain where it will be used.

toolNLP Foundations

What it is

A spaCy pipeline starts with text converted into a Doc containing tokens and can add components for tagging, parsing, lemmatization, entity recognition or custom annotations. Components may depend on attributes produced earlier in the pipeline. Tokenization is a separate initial step with language-specific rules, while trained packages supply learned models and configuration. Rule-based matching can operate on token attributes and combine with statistical predictions. Spans and token offsets link results to source text. Competence includes understanding component order, language and model compatibility, and whether an output is a deterministic rule match or a learned prediction. The library's structured object model does not make all annotations equally reliable.

What the work involves

Choose a language package and enable only components needed for the task, while retaining their prerequisites. Inspect token boundaries, entity spans and other relevant annotations on representative documents. Write precise rules with reviewed positive and negative examples. Batch processing where appropriate and measure end-to-end throughput. When training or tuning, keep documents and related sources separate across splits and evaluate each output type with suitable metrics. Save configuration, rules and model versions together. The useful result is a reproducible pipeline whose structured outputs remain aligned with the source text and whose observed failures are documented for downstream users.

Illustrative example

An illustrative contract-indexing tool uses spaCy to identify organization mentions and match a small set of clause phrases. The engineer checks component order and discovers that one matcher depends on lemmas from a disabled component. They repair the dependency and evaluate full documents held out by template family. Entity spans are shown with their source text, while clause matches are tested against near-miss phrases. Canonical organization linking is handled and measured as a separate stage.

Limits and common mistakes

Language packages and pipeline components have specific coverage and dependencies. Tokenization changes can invalidate span annotations or rules. Statistical entities and parses can be wrong despite a well-formed Doc object, while broad rules can overmatch. Fast processing does not establish task accuracy. spaCy differs from a single language model and from a complete document-understanding application. Pin assets, inspect effective component order and evaluate the exact annotations and downstream decisions required by the application.

Prerequisites

No prerequisites.

Related skills

  • → is an instance of: NLP

Sources and further reading

Last updated: 2026-10-10