Model Retraining
Model retraining updates learned parameters using a new or revised training dataset and procedure. It creates a candidate model whose usefulness must be compared with the deployed version, considering data quality, task changes and regressions before promotion, rather than assuming that fresher data automatically produces a better service.
What it is
Retraining can restart training from an initial state or continue from existing parameters, depending on the model and objective. Triggers include new labels, a scheduled refresh, changed requirements or evidence that deployed performance has deteriorated. Distribution change is a reason to investigate, not sufficient proof that retraining will fix the issue. This differs from changing inference settings, updating retrieval documents or revising prompts. Those interventions may solve a problem without altering weights. The retraining process must preserve the relationship among data, feature definitions, training code and the candidate artifact so comparisons and rollback remain meaningful.
What the work involves
The practitioner diagnoses the failure, validates new data and defines an evaluation that reflects the current task without leaking future information. Candidate comparisons include important subgroups and the serving contract. Useful outputs include a retraining run, artifact lineage and an acceptance decision. Deployment is a separate step with controlled rollout and monitoring. A failed candidate should leave the existing service intact. Retaining a suitable baseline helps distinguish benefits from data refresh, algorithm changes and random variation.
Illustrative example
A demand model begins making poor predictions after a product range changes. Investigation finds that new categories are missing from training and feature preparation. The team updates the data and pipeline, trains a candidate and evaluates it on a time-separated period containing those categories. The candidate is compared with the current model on existing categories as well. It is promoted only if the intended improvement survives those checks and its input contract is compatible with serving.
Limits and common mistakes
Retraining on flawed or delayed labels can make performance worse. A recent dataset may be too small or unrepresentative, and a trigger based on feature drift can miss the real cause of error. Quality requires diagnosis, appropriate validation and a controlled release. Retraining is not a substitute for repairing broken preprocessing or service behavior. An automated schedule should be able to produce no release when the candidate fails acceptance, rather than making deployment an inevitable consequence of completing training.
Prerequisites
Related skills
- → is subcategory of: ML CI/CD
Sources and further reading
- TFX Evaluator component
Documents comparing candidate and baseline models against validation thresholds before downstream model release.
- MLOps automation pipelines
Places retraining triggers, data checks and model validation within a controlled continuous-training pipeline.
Last updated: 2026-10-10