Atlas · skill

ML CI/CD

ML CI/CD extends software delivery pipelines to data processing, training and model validation. It versions the inputs and artifacts needed to produce a model, checks candidate behavior and promotes an accepted version through controlled deployment, so a data-driven update has traceable evidence and a recovery path.

conceptCI/CD & Automation

What it is

A machine-learning release depends on code, data, feature definitions, configuration and trained parameters. ML CI/CD coordinates these artifacts and their checks, rather than treating a model file as an ordinary code build. Continuous training may create candidates when new data or another trigger arrives, but it remains distinct from automatically deploying them. Data validation checks whether training should proceed; model validation checks whether the candidate satisfies performance and compatibility requirements. General CI/CD delivers software changes, while ML CI/CD additionally manages the lineage and behavior of learned artifacts whose outcomes depend on changing input distributions.

What the work involves

The practitioner builds modular pipeline components, pins data and environment references and records artifact lineage. Gates should test data quality, leakage, relevant subgroups and the serving contract. A useful deliverable includes pipeline definitions, candidate evaluation records and a registered model version. Promotion should preserve the exact accepted artifact. Online rollout and monitoring complement offline checks, while rollback includes preprocessing and feature dependencies so the restored model receives inputs consistent with its training and validation.

Illustrative example

A forecasting pipeline receives a refreshed dataset and trains a candidate model. Schema checks catch a changed unit before training; after correction, the candidate is evaluated on time-separated data and critical product groups. The accepted artifact is deployed to a small traffic slice with its preprocessing version. A failed serving-compatibility check blocks promotion even if the forecast error improves, demonstrating that delivery gates cover the whole prediction service.

Limits and common mistakes

Automated retraining can reproduce bad data or silently promote a regression if the gates are weak. Aggregate metrics may conceal failures in important segments. Feature-store or preprocessing changes can break a model without changing its weights. Quality requires lineage, task-relevant validation and controlled promotion. A scheduled pipeline is not continuous improvement by itself: each candidate must earn acceptance under a stable comparison, and the system must handle cases where no new model should be released.

Prerequisites

  • hardGit

    CI/CD is triggered by Git commits and orchestrated around Git branches — Git is the foundation

  • hardDocker

    CI/CD pipelines run in containers and produce container images as artifacts

Related skills

Sources and further reading

Last updated: 2026-10-10