Atlas · skill

Data Versioning

Data versioning identifies and preserves distinct states of datasets and related artifacts over time. It enables reproducible experiments, comparisons and rollback by linking the exact data used to code and configuration, rather than relying on a filename that may point to changing contents.

conceptVersioning & Lineage

What it is

A data version can be an immutable snapshot, a content-addressed artifact or a defined state within a transactional table. Versioning records identity and change; lineage records how an artifact was produced. Reproducibility needs both when transformations depend on upstream sources. Large datasets are often stored outside Git, while lightweight metadata in version control points to their identities. The mechanism differs from backup, whose primary goal is recovery from loss, and from keeping many manually named copies. A useful version must be retrievable and sufficiently described to reconstruct the relevant experiment or analysis.

What the work involves

The practitioner defines what constitutes a dataset release and records content identity, schema and provenance. They connect versions to preparation code, model runs and evaluation outputs, then test retrieval in a clean environment. Useful artifacts include a dataset manifest and policies for storage, access and retention. Historical versions containing sensitive information need the same governance attention as current data. The team also identifies mutable external dependencies and decides whether to snapshot them or document the limits they impose on reproduction.

Illustrative example

Two model experiments report different results against a file called training.csv. Investigation finds that the file changed between runs. The team introduces immutable dataset versions and records their identifiers with each experiment, along with split and preprocessing configuration. A later comparison retrieves both versions and isolates whether the improvement came from the model code or changed examples, instead of treating the shared filename as evidence of identical inputs.

Limits and common mistakes

A version identifier is insufficient if the referenced data has been deleted or access is unavailable. Content hashes detect byte changes but do not explain semantic differences, while snapshots can create substantial storage and privacy obligations. Reproduction may still fail because of nondeterminism or external services. Good versioning states what can be reconstructed and keeps retention choices explicit, rather than promising permanent reproducibility from metadata alone.

Prerequisites

  • hardGit

    DVC extends Git for data versioning — Git is the foundation

Related skills

  • → is subcategory of: Data Engineering
  • ← is an instance of: DVC

Sources and further reading

  • DVC: user guide

    Official data artifact, pipeline and remote-storage versioning documentation.

Last updated: 2026-10-10