Atlas · skill

Experiment Tracking

Experiment tracking records how an experiment was configured, what it produced and how its results were measured. It links runs to data, code and artifacts so practitioners can compare alternatives, locate a selected model and investigate differences, turning scattered execution outputs into interpretable experimental records.

conceptExperiment Tracking & Registry

What it is

A tracking record typically includes parameters, metrics over time, artifacts and identifiers for the execution's inputs. Grouping related runs supports comparisons, while lineage links a saved model to the procedure that created it. This differs from production monitoring, which observes a running service, and from reproducibility, which establishes that a result can be repeated. Tracking supplies evidence for both but does not guarantee either. Metric definitions, evaluation splits and environment conditions determine whether two runs are comparable. A dashboard can organize incompatible measurements just as easily as valid experiments.

What the work involves

The practitioner chooses the information required to interpret the experiment and instruments logging at meaningful points. Dataset and code references should be immutable or otherwise identifiable. Useful outputs include searchable run records and a comparison tied to a specific candidate artifact. Teams use consistent measurement definitions and record failed runs instead of preserving only favorable outcomes. Before selecting a model, they inspect subgroup results and the evaluation protocol so a visible metric improvement is not mistaken for an equivalent comparison.

Illustrative example

A team varies an embedding model and retrieval cutoff while keeping its question set fixed. Each run records index version, query configuration, relevance metrics and example results. The comparison reveals that one apparent improvement used a newer corpus, so it is rerun under matched conditions. The selected configuration remains linked to the resulting index artifact. Later investigation can retrieve the exact run rather than infer its settings from an informal filename.

Limits and common mistakes

Missing inputs or inconsistent metric names make records difficult to interpret. Automatic logging can omit custom preprocessing and hidden state. Recording every artifact indiscriminately can increase storage and expose sensitive examples. Quality requires sufficient relevant metadata, clear comparison conditions and durable artifact links. Tracking should support decisions and scrutiny; the number of logged runs is not evidence of a rigorous experiment, and a well-recorded result can still be affected by leakage or inappropriate evaluation.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

  • MLflow experiment tracking

    Defines runs and the parameters, metrics, artifacts and backend storage used to organize experiment records.

Last updated: 2026-10-10