Model Deployment
Model deployment makes an accepted model artifact available for operational use through a defined interface and environment. It combines packaging, configuration and rollout with readiness, monitoring and rollback, ensuring that the released prediction behavior corresponds to the model and preprocessing that were actually validated.
What it is
Deployment connects a model artifact with runtime dependencies, input transformation, serving code and infrastructure. The result may be an online endpoint, batch job or embedded application. A release can replace traffic immediately, run alongside an existing version or receive a limited share before promotion. This differs from training and from merely copying a model file to a server. Operational correctness requires the whole prediction path to match evaluation conditions. Versioning includes tokenizers, features and preprocessing as well as weights, because a correct artifact can behave differently when paired with incompatible surrounding components.
What the work involves
The practitioner defines the serving contract, packages immutable artifacts and validates the actual deployment environment. Readiness should confirm model availability, while rollout observations cover errors and task outcomes. Useful outputs include a release manifest, deployment record and tested rollback procedure. Access controls and request limits belong in the endpoint. Load tests establish capacity before full traffic. Monitoring should distinguish infrastructure failure from changed prediction quality, and promotion criteria should be specified before a candidate is exposed to users.
Illustrative example
A classifier is released with a new normalization step. The deployment manifest identifies both model and preprocessing versions, and staging checks compare endpoint outputs with the evaluated pipeline. A canary receives limited traffic before promotion. When a regression appears for a known input class, rollback restores the earlier pair together. The exercise verifies that restoring old weights with new preprocessing would not be considered a complete recovery.
Limits and common mistakes
An endpoint can be healthy while returning incorrect predictions. Canary traffic may miss rare cases, and schema compatibility does not establish semantic equivalence. Rollback can be difficult when surrounding data or state has changed. Quality requires artifact identity, tested contracts and observable rollout decisions. Model deployment is complete only when the intended operational path works under its requirements, not when a container starts or a model object loads successfully in a development notebook.
Prerequisites
Related skills
- → is subcategory of: ML CI/CD
- ← is subcategory of: Real-Time Inference
Sources and further reading
- Kubernetes Deployments
Documents controlled application revisions, rollout status and rollback for containerized serving deployments.
- KServe canary rollout strategy
Illustrates model-service revision traffic control and recovery in a supported deployment mode.
Last updated: 2026-10-10