Atlas · skill

Data Cleaning

Data cleaning detects and corrects records that are missing, inconsistent, duplicated, malformed or otherwise unsuitable for an intended use. It combines explicit rules with domain judgment, preserving the difference between a genuine unusual observation and an error that should be repaired or excluded.

Also searchable as: Data Cleansing, Data Cleaning Techniques, data-cleaning, czyszczenie danych

conceptData Quality

What it is

Cleaning transforms data according to assumptions about valid values and relationships. It may standardize formats, reconcile units, remove duplicate representations or handle missing information. These operations have different consequences: dropping a record changes the represented population, while imputing a value adds an estimate rather than recovering the original observation. An outlier may be a valuable rare case, so extremity alone is not evidence of error. Cleaning differs from feature preprocessing, which prepares model inputs, and from curation, which decides broader inclusion and use. The relevant mechanism is a documented correction tied to the task and source evidence.

What the work involves

The practitioner profiles data, identifies failure patterns and creates rules that distinguish valid exceptions from defects. They preserve raw inputs or suitable provenance and record material changes. Useful artifacts include a cleaning specification and before/after checks for affected fields and population coverage. Missing-value treatment should reflect why values are absent, and learned imputation stays within training boundaries when used for modeling. The team samples corrected records and evaluates downstream effects, ensuring that a cleaner-looking dataset has not lost information or introduced systematic distortion.

Illustrative example

A sensor dataset mixes temperature units and contains repeated uploads. The engineer uses source metadata to convert units, identifies duplicate events by stable identifiers and flags implausible readings for review. A rare high temperature is retained because a maintenance report confirms it. Missing readings remain distinguishable from measured zeros, preventing a later analysis from treating an interrupted sensor as evidence that the equipment cooled.

Limits and common mistakes

Automatic repair can hide upstream problems or encode unsupported assumptions. Missingness can carry information, and deduplication can remove legitimate repeated events if identity rules are wrong. Rules learned from the full dataset can leak information into model evaluation. Cleaning quality should be assessed through traceable decisions and downstream validity, not the disappearance of every null or outlier. Some defects should remain flagged until source evidence supports a correction.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10