Atlas · GenAI 2026

Cloud & AI Platform Infrastructure

24 skills · ontology graph below shows relations within this section.

What this domain covers

This edition groups 24 capabilities in Cloud & AI Platform Infrastructure across 6 named categories. The inventory contains 4 concepts and 20 tools. Open an entry for its mechanism, practical workflow, example, limitations, and primary references.

Current category labels: AWS · Azure · Cloud Platforms · Distributed Systems · GCP · Infrastructure as Code

Frequent learning foundations

  1. LLM API Integration supports 2 mapped skills
  2. Python supports 2 mapped skills
  3. Docker supports 1 mapped skill
  4. MLflow supports 1 mapped skill
  5. Shell Scripting supports 1 mapped skill

Skills in this section

Amazon Bedrock
AWS

Amazon Bedrock provides managed access to foundation models and supporting generative-AI services through AWS interfaces. Competence means choosing an appropriate model and invocation path, configuring access and data controls, and operating an application with measured quality, latency and cost rather than relying on the platform's managed-service label.

Amazon SageMaker
AWS

Amazon SageMaker supports AWS data and AI workflows; SageMaker AI is its service for building, training and deploying machine-learning models. The competency is selecting and connecting the appropriate workflow components, controlling their identities and artifacts, and validating the behavior and cost of the resulting training or inference system.

Azure OpenAI Service
Azure

Azure OpenAI provides access to supported OpenAI models through Azure-managed resources and deployment interfaces. Competence involves selecting the model and deployment configuration, applying Azure identity and network controls, and validating application quality and usage while distinguishing this service from the broader Microsoft Foundry platform.

AWS
Cloud Platforms

Amazon Web Services is a cloud platform supplying compute, storage, networking, identity and managed services used by AI systems. The competency is designing a workload from those components with explicit reliability, security and cost decisions, rather than assuming that using a particular provider makes the architecture effective.

Microsoft Azure
Cloud Platforms

Microsoft Azure provides cloud infrastructure, identity, data and managed application services for AI workloads. Competence means choosing and configuring resources around the workload's access, reliability and operational requirements, including clear boundaries between general Azure infrastructure, Azure Machine Learning and Microsoft Foundry services.

Distributed Systems
Distributed Systems

Distributed systems coordinate work across processes or machines that communicate over fallible networks. The competency is designing for partial failure, latency, concurrency and uncertain ordering, so an AI pipeline or service remains correct when one component retries, slows down or loses contact with another.

Google Vertex AI
GCP

Google Vertex AI denotes Google Cloud's managed machine-learning workflows for training, model management and inference. Competence means configuring a reproducible path through those services, governing identities and data access, and validating quality and capacity while keeping individual ML capabilities distinct from broader platform branding.

IaC (Infrastructure as Code)
Infrastructure as Code

Infrastructure as Code manages infrastructure through reviewable configuration or programs rather than undocumented manual changes. The competency is expressing intended resources, controlling state and dependencies, and applying changes safely so environments can be reproduced and differences can be inspected before they affect a running AI workload.

Terraform
Infrastructure as Code

Terraform is an Infrastructure as Code tool that manages resources through provider-backed configuration and tracked state. Competence means understanding what a plan will create, update or replace, controlling state and credentials, and applying changes that respect the lifecycle of data and applications already using those resources.

AWS Fargate
AWS

AWS Fargate supplies managed compute for container workloads without requiring the operator to maintain the underlying server fleet. The competency is defining container resources, networking, identities and service behavior appropriately, including how a task starts, receives work, reports health and stops under failures or updates.

Amazon EMR
AWS

Amazon EMR provides managed environments for distributed data-processing frameworks such as Apache Spark. Competence means configuring jobs and resources around data partitioning, shuffle and storage access, while managing failure recovery and cost so large-scale preparation for AI remains correct and operationally predictable.

Amazon Textract
AWS

Amazon Textract extracts text and structured information from supported document images through managed AWS APIs. Competence means choosing extraction operations, preserving document structure and checking field-level quality, especially when a downstream workflow needs reliable tables, forms or decisions rather than merely recognized characters.

Azure AI Search
Azure

Azure AI Search is a managed retrieval service for indexed and supported connected content, including keyword, vector and hybrid search. The competency is designing content schemas, ingestion and relevance together with access controls so an application retrieves useful evidence that the caller is allowed to see.

Azure Machine Learning
Azure

Azure Machine Learning supports managed development, execution and deployment of machine-learning workflows. Competence means organizing data, environments, jobs and model artifacts into a reproducible process, then configuring inference and operational controls so the deployed model corresponds to the version and preprocessing that were actually evaluated.

Foundry Tools
Azure

Foundry Tools are Microsoft-managed AI services for capabilities such as speech, language, vision and document processing. Competence means integrating the appropriate service for a specific task, interpreting its structured outputs and constraints, and validating quality on real input variations rather than treating the collection as one universal intelligence API.

Microsoft Foundry
Azure

Microsoft Foundry is an Azure platform grouping models, agents and tools with shared management capabilities. Competence means selecting the appropriate components and configuring project access, networking, observation and evaluation, while recognizing that a unified platform interface does not make its individual services behaviorally or operationally identical.

Cloud Platforms
Cloud Platforms

Cloud platforms provide remotely managed compute, storage, networking, identity and higher-level services. The competency is choosing a coherent architecture from these building blocks and managing its operating boundaries, including shared responsibilities, data movement, failure recovery and consumption cost across an AI workload's full lifecycle.

Google Cloud Platform (GCP)
Cloud Platforms

Google Cloud provides infrastructure, data and managed application services used by AI workloads. The competency is composing projects, identities, storage, compute and networking into a reliable architecture, then using workload evidence to choose operational settings and costs rather than assuming that a provider's managed AI services cover the complete system.

Dask
Distributed Systems

Dask schedules parallel computations in Python and offers collections resembling arrays and DataFrames over partitioned data. The competency is constructing useful task graphs, selecting partition sizes and managing memory and communication so scaling a calculation preserves its meaning and avoids spending more effort on coordination than computation.

HPC Cluster Computing
Distributed Systems

HPC cluster computing runs demanding workloads on coordinated compute nodes, commonly through a batch scheduler and shared storage. For AI work, competence means requesting suitable resources, configuring distributed execution and checkpointing, and understanding queueing, interconnect and filesystem behavior so a large job uses its allocation effectively.

Ray
Distributed Systems

Ray is a distributed Python runtime with task and actor abstractions and libraries for training, tuning and serving. Competence means expressing useful parallel work, managing object movement and resource requirements, and handling failures so a distributed AI application remains understandable and efficient as execution spans multiple workers.

Cloud Run
GCP

Cloud Run is Google Cloud's managed platform for running application code and containers through services, jobs and other supported execution modes. Competence means choosing the right mode and configuring concurrency, identity, resources and startup behavior so an AI workload fits the platform's lifecycle and capacity constraints.

Google Cloud Build
GCP

Google Cloud Build executes configured build steps to test source and produce deployable artifacts on Google Cloud. Competence means defining a reproducible build, controlling its identity and triggers, and preserving the relationship between reviewed source, validation evidence and the container or package eventually deployed.

Google Cloud Data Fusion
GCP

Cloud Data Fusion is Google Cloud's managed data-integration service for designing and running pipelines through a visual interface and supported connectors. Competence means specifying schemas, transformations and execution settings precisely, then validating data movement, recovery and cost beyond the appearance of a successfully connected pipeline diagram.