Data Drift
Data drift is a change in the distribution of data observed by a system compared with a reference period or dataset. It can signal that deployment inputs no longer resemble development data, but a detected change does not by itself prove model degradation or explain whether retraining is the right response.
What it is
Drift can concern features, missingness, categories, text characteristics or predictions, depending on what is compared. Statistical tests, distances or discriminators can detect different forms of change. The reference and current windows define the comparison, while sample size and threshold affect sensitivity. Data drift differs from concept drift, where the relationship between inputs and the target changes. An input distribution can shift without harming a model, and target relationships can change even when simple feature summaries look stable. The distinction prevents treating every distribution alert as direct evidence of accuracy loss.
What the work involves
The practitioner selects meaningful reference and current windows, validates data types and monitors missing or newly introduced values separately where needed. It chooses tests suitable for the variable and checks false alarms under expected seasonal variation. Useful artifacts include drift reports, thresholds and an investigation procedure linked to quality or outcome measurements. Alerts should identify affected features and populations. Before retraining, the team checks for pipeline defects, changes in users or genuine behavior shifts, then assesses whether these changes affect the task the model performs.
Illustrative example
A demand model begins receiving more records from a new region. Monitoring detects a changed distribution in region and order size. The team checks whether the ingestion pipeline is correct and examines forecast errors for the new population once outcomes arrive. If quality remains acceptable, retraining may be unnecessary. If missing values increased because a source field was renamed, fixing the pipeline is more appropriate than teaching the model to accommodate corrupted input.
Limits and common mistakes
Large samples can flag small harmless differences, while small samples may miss important shifts. Marginal feature tests can also miss changed joint relationships. Alert thresholds require context and cannot be copied as universal standards. Drift should be investigated with data quality and task performance, avoiding automatic retraining based on one statistic. Its value lies in identifying a change worth examining, with explicit uncertainty about the effect on model behavior.
Prerequisites
Related skills
- → is part of: ML Monitoring
Sources and further reading
- Evidently data drift explainer
Defines reference/current distribution comparisons and documents test-selection and missingness considerations.
Last updated: 2026-10-10