Atlas · GenAI 2026

LLMOps, Model Serving & Inference Optimization

40 skills · ontology graph below shows relations within this section.

What this domain covers

This edition groups 40 capabilities in LLMOps, Model Serving & Inference Optimization across 14 named categories. The inventory contains 18 concepts and 22 tools. Open an entry for its mechanism, practical workflow, example, limitations, and primary references.

Current category labels: API Gateways & Routing · CI/CD & Automation · Containerization · Cost & FinOps · Deployment Infrastructure · Experiment Tracking & Registry · GPU & Kernels · Inference Optimization · and 6 more

Frequent learning foundations

  1. Docker supports 4 mapped skills
  2. LLM Inference Serving supports 4 mapped skills
  3. Transformer Architecture supports 4 mapped skills
  4. Inference Optimization supports 3 mapped skills
  5. Git supports 2 mapped skills

Skills in this section

LLM API Gateway
API Gateways & Routing

An LLM API gateway sits between applications and model providers to apply common access, routing and operational policies. It can manage credentials, budgets, rate limits and fallback paths, giving practitioners a controlled entry point while preserving the provider-specific behavior needed by each application.

LiteLLM
API Gateways & Routing

LiteLLM provides a Python interface and proxy gateway for calling models across providers through a common application contract. Practitioners configure model aliases, credentials, routing and budgets, then test the features their applications depend on instead of assuming that a normalized API makes every model interchangeable.

CI/CD
CI/CD & Automation

CI/CD automates the integration, validation and delivery of software changes through repeatable pipelines. For AI applications, it connects code and configuration changes to appropriate checks and controlled releases, producing a traceable deployable artifact and a deliberate decision about when that artifact reaches an environment.

ML CI/CD
CI/CD & Automation

ML CI/CD extends software delivery pipelines to data processing, training and model validation. It versions the inputs and artifacts needed to produce a model, checks candidate behavior and promotes an accepted version through controlled deployment, so a data-driven update has traceable evidence and a recovery path.

Docker
Containerization

Docker packages software and its runtime dependencies into container images and runs them as isolated processes. In AI work, practitioners use it to create repeatable training or serving environments, manage filesystem and network configuration and make the exact application package identifiable across development and deployment.

AI Cost Optimization
Cost & FinOps

AI cost optimization changes architecture or execution choices to reduce the resources needed for an acceptable task outcome. It considers model selection, context, caching, batching and unnecessary calls together, measuring savings against quality and latency so a cheaper request does not create more expensive failure or rework.

AI FinOps
Cost & FinOps

AI FinOps applies financial-management practices to AI consumption so teams can understand spending, assign ownership and connect resources to useful outcomes. It combines usage attribution, forecasting and optimization decisions across engineering, finance and product, accounting for both variable model consumption and the infrastructure supporting AI workloads.

Semantic Caching
Cost & FinOps

Semantic caching reuses a stored model response when a new request is judged meaningfully equivalent to a previous one. It typically retrieves candidate requests through embeddings and applies a similarity rule, allowing reuse beyond exact text matches while requiring safeguards for freshness, context and user-specific information.

Serverless AI
Deployment Infrastructure

Serverless AI runs model-related work on infrastructure whose worker lifecycle and scaling are managed by a platform. Practitioners package the workload, select resources and tune initialization and concurrency so variable demand can be served without manually operating a fixed fleet, while accounting for startup delay and platform limits.

MLflow
Experiment Tracking & Registry

MLflow is a platform for recording experiments and managing model-related artifacts and lifecycle information. Practitioners log configurations, metrics and outputs, identify candidate versions and connect evaluation to release decisions, making experimental results discoverable and traceable without assuming that tracking alone makes a model suitable for deployment.

Weights & Biases
Experiment Tracking & Registry

Weights & Biases is a platform for recording and comparing machine-learning experiments and their related artifacts. Practitioners instrument runs, inspect measurements and preserve configuration and data references, making collaborative analysis easier while ensuring that the tracked evidence is sufficient to interpret or reproduce a result.

GPU Kernel Programming
GPU & Kernels

GPU kernel programming implements numerical operations as parallel work executed on a graphics processor. Practitioners choose thread and memory layouts, combine operations and handle numerical and shape constraints, aiming to reduce actual execution time while preserving the required computation across representative inputs and hardware.

Inference Optimization
Inference Optimization

Inference optimization reduces the time or resources required to run a trained model while preserving acceptable output behavior. It combines workload measurement with choices such as batching, precision, kernels and memory management, balancing throughput and request latency under the actual input lengths and concurrency of the application.

KV Cache Optimization
Inference Optimization

KV cache optimization manages the attention keys and values retained during autoregressive generation. It reduces memory waste or storage requirements and controls reuse across eligible prefixes, allowing a server to handle context and concurrency efficiently while checking whether any compression changes model output quality.

Speculative Decoding
Inference Optimization

Speculative decoding accelerates autoregressive generation by proposing several tokens with a cheaper draft process and checking them with the target model. A verification procedure accepts valid proposals and corrects rejected ones, reducing sequential target-model work when the draft agrees often enough to justify its overhead.

Ollama
Local Inference Runtimes

Ollama provides a model-running service and interfaces for using supported language models, commonly on local hardware. Practitioners select and manage model artifacts, configure context and memory behavior and connect applications through its API, verifying whether the selected execution path is local and suitable for the available resources.

llama.cpp
Local Inference Runtimes

llama.cpp is a C/C++ inference project for running supported language models across CPUs and accelerator backends. Practitioners prepare compatible model artifacts, choose quantization and hardware placement and test generation settings, balancing memory footprint and speed against the quality required by the application.

Kubernetes
Orchestration

Kubernetes orchestrates containerized workloads through declarative resources and controllers that maintain desired state. For AI systems, practitioners configure scheduling, resource requests, service access and lifecycle behavior so training or inference containers run reliably within cluster capacity, including accelerator and model-loading constraints.

Kubeflow
Pipeline Orchestration

Kubeflow is a Kubernetes-oriented ecosystem for machine-learning development and workflows. Practitioners use its components to organize experiments, training and pipeline execution on a cluster, preserving artifacts and dependencies while managing the operational requirements inherited from Kubernetes and the selected ML components.

BentoML
Serving Runtimes

BentoML is a Python framework for building model inference services and packaging their dependencies for deployment. Practitioners define API contracts, initialize models and configure serving resources or batching, turning an inference script into an operational service whose inputs, outputs and lifecycle can be tested.

KServe
Serving Runtimes

KServe is a Kubernetes-based platform for deploying predictive and generative model inference services. Practitioners declare model resources and serving runtimes, configure scaling and networking and choose supported rollout behavior, giving model endpoints a managed lifecycle while keeping model quality and application access policy separate.

LLM Inference Serving
Serving Runtimes

LLM inference serving operates language models as services that accept requests and generate outputs under concurrent demand. It combines model execution with tokenization, scheduling, streaming and resource management, allowing practitioners to meet latency and capacity requirements while maintaining a stable request contract and controlled access.

Ray Serve
Serving Runtimes

Ray Serve is a distributed serving framework for composing and operating model-backed services. Its LLM facilities place inference engines within scalable deployments and expose supported APIs, helping practitioners manage replicas, routing and multi-node execution while preserving engine-specific configuration and application behavior.

SGLang
Serving Runtimes

SGLang is a serving and execution framework for language and multimodal models, with techniques for reusing prefixes and managing structured generation. Practitioners configure the supported model runtime, memory and scheduling behavior and verify API and output requirements under the application's actual workload.

vLLM
Serving Runtimes

vLLM is an inference and serving engine for supported language models, designed to manage concurrent generation and model memory efficiently. Practitioners configure model loading, scheduling, precision and API behavior, then test performance and output quality using the input lengths and request patterns expected in deployment.

Model Retraining
CI/CD & Automation

Model retraining updates learned parameters using a new or revised training dataset and procedure. It creates a candidate model whose usefulness must be compared with the deployed version, considering data quality, task changes and regressions before promotion, rather than assuming that fresher data automatically produces a better service.

Reproducibility
CI/CD & Automation

Reproducibility makes a result repeatable or independently checkable by preserving the inputs, environment and procedure that produced it. In AI work, the practitioner versions data and code, records model and hardware conditions and controls randomness, while stating the tolerance or behavioral equivalence expected from a repeat execution.

Model Deployment
Deployment Infrastructure

Model deployment makes an accepted model artifact available for operational use through a defined interface and environment. It combines packaging, configuration and rollout with readiness, monitoring and rollback, ensuring that the released prediction behavior corresponds to the model and preprocessing that were actually validated.

Experiment Tracking
Experiment Tracking & Registry

Experiment tracking records how an experiment was configured, what it produced and how its results were measured. It links runs to data, code and artifacts so practitioners can compare alternatives, locate a selected model and investigate differences, turning scattered execution outputs into interpretable experimental records.

CUDA
GPU & Kernels

CUDA is NVIDIA's parallel-computing platform and programming environment for executing work on supported GPUs. Practitioners use its memory, kernel and execution model directly or through libraries, understanding compatibility and synchronization so accelerated AI operations are both correct and effective within the application's performance constraints.

GPU Acceleration
GPU & Kernels

GPU acceleration moves suitable computation onto a graphics processor to exploit parallel execution and high memory bandwidth. Practitioners identify work that can benefit, manage data placement and precision and measure the whole application, ensuring that transfer and scheduling overhead do not outweigh faster numerical operations.

FlashAttention
Inference Optimization

FlashAttention is an exact attention implementation that reduces transfers between GPU memory levels through tiling and fused computation. It computes attention without materializing the full intermediate attention matrix in high-bandwidth memory, improving the execution and memory profile where the supported shapes and hardware benefit from that strategy.

OpenVINO
Inference Optimization

OpenVINO is a toolkit for converting, optimizing and running model inference on supported devices, including Intel CPU, GPU and NPU paths. Practitioners prepare compatible representations, select runtime devices and performance settings and validate predictions, managing the differences between conversion success, device support and useful application performance.

TensorRT
Inference Optimization

TensorRT is NVIDIA's inference SDK for turning supported trained models into optimized GPU execution engines and running those engines. Practitioners define shape and precision requirements, build and validate the artifact and integrate the runtime, checking hardware compatibility and output quality before treating compilation as a successful deployment.

ONNX
Model Interchange & Portability

ONNX is an open representation for machine-learning models that describes computation as a graph of typed operations and parameters. Practitioners export and inspect models in this format to move them between compatible tools, preserving input semantics and checking operator versions rather than assuming export guarantees portable behavior.

ONNX Runtime
Model Interchange & Portability

ONNX Runtime executes supported model graphs using CPU and accelerator implementations selected through execution providers. Practitioners load the model, configure providers and session behavior and inspect actual operation placement, validating that the deployed execution preserves predictions and meets performance requirements on the target hardware.

NVIDIA Triton Inference Server
Serving Runtimes

NVIDIA Triton Inference Server operates model endpoints through configurable backends, scheduling and version management. Practitioners define input and output contracts, batching and instance placement, connecting inference runtimes to a server that can handle concurrent requests and expose operational behavior across different model types.

TensorRT-LLM
Serving Runtimes

TensorRT-LLM is NVIDIA's library and runtime ecosystem for optimized language-model inference on supported GPUs. Practitioners configure model execution, attention memory, batching and parallelism, using the documented backend for their version and validating both output quality and workload performance rather than treating the product name as a fixed compilation workflow.

TorchServe
Serving Runtimes

TorchServe is a model server for PyTorch models with packaging, handlers, workers and request-serving configuration. The skill remains relevant to maintaining existing deployments and understanding their contracts. Its official documentation states that the project is no longer actively maintained and has no planned fixes or security patches.

Real-Time Inference
Inference Serving

Real-Time Inference delivers model outputs within the response constraints of an interactive or time-sensitive application. Practitioners budget queueing, preprocessing, execution and delivery together, controlling load and failure behavior so the service meets its stated latency requirement while preserving prediction quality under concurrent demand.