Cloud Platforms
Cloud platforms provide remotely managed compute, storage, networking, identity and higher-level services. The competency is choosing a coherent architecture from these building blocks and managing its operating boundaries, including shared responsibilities, data movement, failure recovery and consumption cost across an AI workload's full lifecycle.
What it is
A cloud platform offers programmable resources with service-specific limits and pricing. Virtual machines provide relatively direct control, containers package applications and managed services take over selected operational tasks. Regions, availability boundaries and resource identities shape placement and access. These concepts generalize across providers, but equivalent-looking services can have different semantics and guarantees. Cloud-platform competence is broader than an individual model API or framework. AI workloads combine data preparation, training, serving and observation, so their architecture must consider both persistent data and intermittent or continuously running computation.
What the work involves
The practitioner inventories workload requirements, compares service options and documents why a deployment needs particular capacity, locations and connectivity. They define access, data retention, observability and recovery, then estimate cost from expected usage and test assumptions with a representative workload. They use reproducible configuration and identify who operates each boundary. Useful work yields an architecture with clear trade-offs and practical evidence, including a way to stop unused resources, recover failed operations and detect when workload growth exceeds the original design.
Illustrative example
A team chooses where to run nightly model retraining and daytime prediction. The engineer compares batch compute with a persistent serving endpoint, includes data-transfer and storage costs and keeps training credentials separate from inference credentials. A staging exercise checks that an interrupted training run cannot overwrite the last approved model. The architecture record explains capacity choices and when demand would justify revisiting them.
Limits and common mistakes
Provider abstractions do not erase quotas, network latency or the operator's responsibilities. Multi-cloud deployments can add coordination costs rather than automatically increasing reliability. A low unit price may still produce a costly architecture through idle capacity or repeated transfer. Check actual service semantics, access boundaries and total lifecycle cost. The appropriate platform depends on the workload; cloud competence is reasoned selection and operation, not allegiance to a vendor.
Prerequisites
Related skills
- ← is an instance of: Google Cloud Platform (GCP)
- → is subcategory of: Distributed Systems
Sources and further reading
- Google Cloud Well-Architected Framework
Supports general workload architecture, operational, security, reliability and cost decisions.
- AWS Well-Architected Framework
Provides an independent provider framework for assessing cloud workload trade-offs.
Last updated: 2026-10-10