Atlas · skill

DVC

DVC manages versioned data and model artifacts alongside version-controlled project metadata. It links large files stored outside Git to their recorded identities and can describe pipeline dependencies, helping teams reproduce experiments and compare results without committing every dataset byte to the source repository.

toolVersioning & Lineage

What it is

DVC stores lightweight tracking metadata in the project while keeping actual data in a cache and optional remote storage. Content identity lets the tooling determine which artifact a project version references and retrieve it when available. Pipeline stages can declare dependencies, commands and outputs, making the data-processing graph explicit. This complements Git rather than replacing it: code and metadata remain versioned together, while large artifacts use suitable storage. A DVC pointer is therefore a reference to data, not the data itself, and reproducibility depends on remote availability, permissions and a sufficiently specified execution environment.

What the work involves

The practitioner tracks important datasets and model artifacts, configures an approved remote and commits their metadata with the relevant code changes. They define pipeline stages where useful and test checkout or reproduction in a clean workspace. Useful artifacts include a versioned pipeline specification and storage-access policy. The team verifies which files are tracked, which remain unrecorded and how sensitive historical versions are retained. Cache cleanup and remote retention need coordination so a valid project commit does not later point to artifacts that can no longer be obtained.

Illustrative example

A model experiment uses a prepared dataset too large for Git. The engineer tracks it with DVC, uploads it to restricted remote storage and commits the metadata alongside the training configuration. A colleague checks out the commit and retrieves the exact artifact, then reproduces the evaluation. When another dataset version improves results, the comparison can distinguish changes in examples from changes in code because both versions remain explicitly identified.

Limits and common mistakes

DVC cannot reproduce data that was never uploaded or has been deleted from its remote. Content identity does not explain semantic changes, and pipeline declarations cannot remove nondeterminism from arbitrary commands. Large historical versions also create storage and privacy obligations. The practitioner should test actual retrieval and reconstruction, document environment dependencies and avoid equating a committed metadata file with guaranteed long-term access to the referenced artifact.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

  • DVC: user guide

    Official data artifact, pipeline and remote-storage versioning documentation.

Last updated: 2026-10-10