Amazon SageMaker
Amazon SageMaker supports AWS data and AI workflows; SageMaker AI is its service for building, training and deploying machine-learning models. The competency is selecting and connecting the appropriate workflow components, controlling their identities and artifacts, and validating the behavior and cost of the resulting training or inference system.
What it is
The canonical Atlas label covers a name with more than one current scope. AWS documentation distinguishes the broader SageMaker data, analytics and AI platform from SageMaker AI, the managed machine-learning service historically called SageMaker. Training jobs, model artifacts and hosted inference are central to the latter's role. Unlike Bedrock's access to supported hosted foundation models, these workflows support substantial control over training code, frameworks and model deployment. Managed resources reduce server administration but do not choose a valid dataset, estimator, evaluation design or resource configuration for the team.
What the work involves
A practitioner defines the data and container inputs for a job, chooses compute suited to the model and sets execution roles with appropriate resource access. They preserve the relationship between training configuration, resulting artifact and evaluation evidence. Before deployment, they choose an inference mode matching traffic and latency requirements and test serialization, preprocessing and capacity. They monitor failures and spending, retiring unused resources. Useful work produces a repeatable path from approved training inputs to a validated serving artifact with clear operational ownership.
Illustrative example
An engineer trains a shipment-delay classifier using a custom container and a versioned dataset. The job writes a model artifact and evaluation report to designated storage. The engineer deploys the approved artifact to an inference endpoint, tests raw request preprocessing and checks that the endpoint's role cannot access the training team's unrelated datasets. A staging traffic exercise informs instance sizing and the release plan.
Limits and common mistakes
The broader SageMaker name should not imply that every component is part of one identical ML service. A successful training job does not establish valid predictions, and endpoint capacity or idle notebooks can continue to generate cost. Permission, network and artifact configuration remain operator decisions. Check train-serving consistency, dependency versions and resource teardown. Platform metadata helps trace work but cannot repair leakage or an inappropriate evaluation design.
Prerequisites
Related skills
- → is an instance of: MLOps
- → is part of: AWS
Sources and further reading
- What is Amazon SageMaker AI?
Explicitly distinguishes SageMaker AI from the broader SageMaker platform and documents training and deployment scope.
Last updated: 2026-10-10