Atlas · skill

Kubernetes

Kubernetes orchestrates containerized workloads through declarative resources and controllers that maintain desired state. For AI systems, practitioners configure scheduling, resource requests, service access and lifecycle behavior so training or inference containers run reliably within cluster capacity, including accelerator and model-loading constraints.

toolOrchestration

What it is

Users declare resources such as Pods, Deployments, Jobs and Services, and controllers reconcile actual cluster state with those declarations. The scheduler places workloads on eligible nodes according to resource and policy constraints. Kubernetes manages the application lifecycle, not the model's numerical execution. Accelerator workloads require drivers and device integration, while persistent model storage and networking require additional configuration. This differs from Docker's packaging and local container execution. A container that starts successfully may still be unready to serve because model loading or a dependency is incomplete, so lifecycle probes must represent actual service state.

What the work involves

The practitioner specifies resource requests and limits, node eligibility and the appropriate workload controller. Readiness checks should reflect model availability, and graceful termination should account for active requests. Useful outputs include deployment manifests, scaling policy and operational runbooks. Tests cover node loss, startup failure and overload. GPU scheduling, storage transfer and initialization time need realistic capacity planning. Autoscaling signals should match the service's bottleneck rather than assuming CPU use is sufficient for an accelerator-bound model.

Illustrative example

An inference service runs on GPU nodes with model weights supplied from versioned storage. Its readiness probe remains false until the model and tokenizer are loaded. A rolling update creates new replicas before withdrawing old ones, with termination handling for active streams. A test removes a node and inspects whether replacement capacity becomes ready within the service's requirements. The team also checks that a burst does not schedule more replicas than eligible accelerators can support.

Limits and common mistakes

Kubernetes cannot create unavailable accelerator capacity or guarantee model quality. Poor probe settings can cause restart loops during slow initialization, and autoscaling can lag behind bursts. Cluster abstraction does not remove driver and hardware differences. Quality requires tested lifecycle and capacity behavior, controlled access and observability. Kubernetes adds operational complexity, so its use should solve concrete scheduling or reliability needs rather than be treated as a prerequisite for every model deployment.

Prerequisites

  • hardDocker

    Kubernetes orchestrates containers — you must understand what a container is before orchestrating thousands of them

Related skills

  • → is an instance of: Container Orchestration

Sources and further reading

  • Kubernetes overview

    Explains declarative container orchestration, desired state and cluster workload management.

  • Kubernetes GPU scheduling

    Documents device-plugin and resource requirements for scheduling accelerator workloads.

Last updated: 2026-10-10