Atlas · skill

Data Curation

Data curation selects, organizes and documents data for a defined purpose. It combines relevance assessment, provenance, deduplication and quality review so a dataset represents intentional choices about coverage and use rather than merely the accumulation of all available records.

conceptDataset Curation

What it is

Curation begins with the intended use and asks which examples belong, which should be excluded and what evidence accompanies those decisions. It may include cleaning and labeling, but its scope is broader: selection criteria determine the population and content that a model or analysis will encounter. Provenance records origin and transformations, while documentation explains composition and known gaps. Curation also considers rights, sensitive information and downstream restrictions. The mechanism is controlled selection with traceable decisions, not simply improving file formatting or maximizing dataset size. Removing low-quality records can help while also changing whose cases remain represented.

What the work involves

A practitioner defines inclusion criteria, inspects source characteristics and implements repeatable filtering and deduplication. They review samples around filter thresholds and record exclusion reasons. Useful artifacts include a dataset manifest, provenance records and a coverage report comparing the selected data with the intended use. When human judgment is involved, guidelines and disagreement handling keep decisions consistent. The process preserves versions so consumers can understand changes, and checks whether curation removes difficult but important cases that evaluation or training still needs to cover.

Illustrative example

A team builds a retrieval corpus from technical documentation. It keeps authoritative manuals, separates obsolete versions and removes duplicate navigation pages, but retains rare troubleshooting sections that simple length filters would discard. Each document records origin and applicability. A coverage review checks whether the curated corpus supports common tasks and unusual failures, revealing that a smaller, deliberate collection can be more useful than an unfiltered crawl.

Limits and common mistakes

Selection introduces bias, and curation rules can quietly erase minority cases or inconvenient evidence. Deduplication may remove legitimate repeated events, while provenance can be incomplete for third-party data. A curated dataset still needs task-specific evaluation and does not inherit truth or permission from its packaging. Quality is demonstrated by documented choices and representative coverage, including known exclusions, rather than a broad claim that the data is clean.

Prerequisites

  • Data curation is a specialized ETL pipeline — general pipeline design skills are the foundation

  • PII removal is a core step in curation pipelines — understanding what constitutes PII and how to mask it is required

Related skills

Sources and further reading

  • Datasheets for Datasets

    Primary proposal for documenting dataset motivation, composition, collection and intended uses.

Last updated: 2026-10-10