Atlas · GenAI 2026

Data Engineering & Pipelines

43 skills · ontology graph below shows relations within this section.

What this domain covers

This edition groups 43 capabilities in Data Engineering & Pipelines across 21 named categories. The inventory contains 25 concepts and 18 tools. Open an entry for its mechanism, practical workflow, example, limitations, and primary references.

Current category labels: Data Architecture · Data Governance & Catalog · Data Governance & Contracts · Data Ingestion · Data Modeling & Design · Data Privacy & Compliance · Data Quality & Contracts · Data Quality & Integration · and 13 more

Frequent learning foundations

  1. SQL supports 4 mapped skills
  2. Data Curation supports 3 mapped skills
  3. Python supports 3 mapped skills
  4. PII Management supports 2 mapped skills
  5. Apache Kafka supports 1 mapped skill

Skills in this section

Data Mesh
Data Architecture

Data mesh is an organizational and architectural approach in which domain teams own data products for other teams to consume. It combines distributed ownership with shared infrastructure and federated governance, addressing the coordination problem of supplying trustworthy analytical data across a large organization.

Databricks Unity Catalog
Data Governance & Catalog

Databricks Unity Catalog is a governance layer for organizing and controlling access to data and AI assets in Databricks. The skill involves configuring its object hierarchy, privileges and evidence of asset use so shared analytics and model workflows operate under an explicit permission model.

Data Contracts
Data Governance & Contracts

Data contracts specify what a producer promises about data and what consumers can rely on. They make schema, semantics, ownership, quality and change expectations explicit, allowing pipelines and teams to detect incompatible changes before those changes silently alter analysis, model inputs or operational decisions.

Document Parsing
Data Ingestion

Document parsing converts files such as PDFs, HTML and office documents into structured content that downstream systems can use. It preserves meaningful elements and provenance, including headings, tables and page locations, rather than treating every document as a single undifferentiated text string.

Web Scraping
Data Ingestion

Web scraping extracts structured information from web pages through repeatable collection and parsing. The skill combines discovery, fetching, page interpretation and data validation so changing web content can become a traceable dataset without confusing presentation artifacts, duplicate pages or incomplete loads with reliable source records.

Data Modeling
Data Modeling & Design

Data modeling designs how entities, relationships and measurements are represented in stored data. It connects domain meaning to schemas, keys and constraints, helping applications and analytical systems preserve valid relationships while choosing storage structures that support their expected queries and changes.

PII Management
Data Privacy & Compliance

PII management governs personal information across collection, storage, use, sharing and deletion. It combines data inventory, purpose and access decisions with technical controls, recognizing that privacy risks extend beyond obvious identifiers to combinations of attributes, derived records and copies created by analytics or AI systems.

Data Quality Management
Data Quality & Contracts

Data quality management defines and maintains the properties data needs for a particular use. It turns expectations about correctness, completeness, consistency and timeliness into checks and remediation processes, so problems are investigated at their source rather than repeatedly patched in downstream analysis or model training.

Data Observability
Data Quality & Integration

Data observability uses metadata, measurements and lineage to understand the health of data pipelines and their outputs. It helps detect unexpected changes in freshness, volume, schema or distributions and trace their downstream effects, shortening the path from an unreliable result to an actionable explanation.

Entity Resolution
Data Quality & Integration

Entity resolution determines which records refer to the same real-world entity across imperfect or disconnected sources. It combines identity rules, similarity evidence and uncertainty handling to link or consolidate records without assuming that matching names are unique or that different identifiers always imply different entities.

NoSQL
Databases & Storage

NoSQL describes database families that use storage and access models beyond the traditional relational-table interface. Document, key-value, wide-column and graph databases offer different ways to represent and retrieve data, so competence means choosing a specific model from access patterns and consistency needs rather than treating NoSQL as one technology.

Data Curation
Dataset Curation

Data curation selects, organizes and documents data for a defined purpose. It combines relevance assessment, provenance, deduplication and quality review so a dataset represents intentional choices about coverage and use rather than merely the accumulation of all available records.

Data Labeling & Annotation
Dataset Curation

Data labeling and annotation turn observations into supervised examples or evaluation judgments using a defined task and guidance. The skill designs labels, manages annotation work and checks agreement and errors, recognizing that a label is a measurement produced by a process rather than unquestionable ground truth.

Dataset Engineering
Dataset Curation

Dataset engineering designs and maintains datasets that support valid model development. It controls sampling, schemas, splits, transformations and provenance so training, validation and testing reflect the intended task, with particular attention to leakage and the difference between available data and data available at prediction time.

Evaluation Data Engineering
Dataset Curation

Evaluation data engineering creates and maintains examples that measure an AI system's intended behavior. It defines coverage, reference judgments and versioned test conditions, allowing teams to compare changes and diagnose failures without mistaking a convenient sample or familiar benchmark for evidence about their actual deployment.

Training Data Curation
Dataset Curation

Training data curation selects and prepares examples that teach a model the intended task or behavior. For supervised fine-tuning, it aligns inputs, target responses and training loss with the desired deployment, while controlling duplicates, inconsistent instructions and examples that teach behavior the application should not reproduce.

Apache Spark
Distributed Processing

Apache Spark is a distributed engine for processing data through coordinated work across multiple executors. The skill involves expressing transformations, understanding partitioning and execution plans, and diagnosing the movement of data so large batch or streaming workloads run correctly and efficiently.

ETL Pipeline Design
ETL/ELT

ETL pipeline design organizes how data is extracted, transformed and loaded into a target system. It defines dependencies, data semantics and recovery behavior, choosing whether transformations occur before or after loading so downstream analytics and AI workloads receive dependable, traceable inputs.

Apache Kafka
Streaming

Apache Kafka is an event-streaming platform that stores records in partitioned logs and lets consumers process them independently. The skill designs topics, keys and consumer behavior so events can support decoupled services, replay and data pipelines with explicit ordering, retention and failure assumptions.

Stream Processing
Streaming

Stream processing computes results from events that arrive continuously rather than from a fixed, completed dataset. It manages time, state and incomplete information, allowing systems to produce ongoing aggregates or decisions while accounting for late events, out-of-order delivery and recovery from failures.

Event-Driven Architecture
Streaming & Messaging

Event-driven architecture connects components through records of things that happened, allowing producers and consumers to evolve with less direct coordination. The skill designs event meaning, delivery and handling so decoupling does not turn into unclear ownership, inconsistent state or repeated side effects.

Apache Iceberg
Table Formats & Storage

Apache Iceberg is an open table format for large analytical datasets stored as files. It adds table metadata, snapshots and evolution mechanisms above the file layer, allowing compatible engines to read consistent table states without treating a directory listing as the complete definition of a table.

dbt
Transformation

dbt organizes data transformations as versioned projects with declared dependencies, tests and documentation. The skill uses those projects to build dependable analytical models in a supported data platform, making transformation logic reviewable and connecting source tables to the datasets consumed by analysts and AI pipelines.

Data Versioning
Versioning & Lineage

Data versioning identifies and preserves distinct states of datasets and related artifacts over time. It enables reproducible experiments, comparisons and rollback by linking the exact data used to code and configuration, rather than relying on a filename that may point to changing contents.

BigQuery
Warehouses & Lakehouses

BigQuery is Google Cloud's managed analytical data warehouse for querying large datasets with SQL and related services. Competence includes modeling tables, controlling scanned work, configuring access and interpreting execution behavior so analyses and AI data preparation are correct, repeatable and economical for the actual workload.

Databricks
Warehouses & Lakehouses

Databricks is a platform for data engineering, analytics and AI workflows built around shared data and compute services. The skill connects ingestion, transformation, experimentation and deployment within the platform while choosing appropriate governance and operational practices for the specific workload.

Snowflake
Warehouses & Lakehouses

Snowflake is a managed data platform whose analytical architecture separates persistent storage from virtual warehouses that execute queries. The skill designs data structures, access and compute use so teams can run dependable analytics and AI-related data workflows while understanding isolation, concurrency and consumption.

Apache Airflow
Workflow Orchestration

Apache Airflow orchestrates workflows expressed as directed acyclic graphs of tasks. It schedules and supervises dependent work, records execution state and supports recovery, allowing data and model pipelines to run repeatedly while making their operational history and dependency structure visible.

BeautifulSoup
Data Ingestion

BeautifulSoup is a Python library for navigating and extracting information from HTML and XML parse trees. The skill uses structured selection and text handling to turn downloaded markup into records, while distinguishing the parser's view of a document from the content a browser may render dynamically.

Data Ingestion
Data Ingestion

Data ingestion brings records from source systems into a destination where they can be processed or consumed. It defines how delivery, schema changes and progress are tracked, so collecting data includes reliable replay and correction behavior rather than merely moving bytes between services.

Optical Character Recognition (OCR)
Data Ingestion

Optical character recognition converts images of writing into machine-readable text. The skill selects and evaluates recognition pipelines for the actual document conditions, including language, scan quality and layout, so extracted characters can support search or analysis without hiding uncertainty behind apparently clean text.

Scrapy
Data Ingestion

Scrapy is an asynchronous Python framework for crawling websites and extracting structured records. It coordinates requests, responses, spiders and item pipelines, giving practitioners a repeatable collection architecture with explicit handling for scheduling, retries, normalization and the operational behavior of a crawl.

Tesseract
Data Ingestion

Tesseract is an open-source OCR engine for recognizing text in images. The skill configures language resources, segmentation and preprocessing for a document collection, then measures recognition errors and preserves source context so the engine's output can be used responsibly in a larger extraction workflow.

CVAT
Dataset Curation

CVAT is an annotation platform for image and video datasets. The skill configures tasks, labels and review workflows, using spatial and temporal annotation tools to produce consistent training or evaluation targets while preserving the distinction between efficient annotation and accurate interpretation of the source material.

Data Augmentation
Dataset Curation

Data augmentation creates additional training views by applying transformations that preserve the relevant target meaning. It encourages a model to tolerate expected variation, such as image changes or input noise, while requiring careful judgment about which transformations remain valid for the task and label.

Label Studio
Dataset Curation

Label Studio is a configurable platform for annotating data across modalities such as text, images and audio. The skill builds annotation interfaces and review workflows that match a task's measurement needs, integrating model assistance where useful while checking the quality and provenance of exported labels.

Kedro
ETL/ELT

Kedro is a Python framework for structuring data-science and machine-learning projects as explicit pipelines. It separates processing functions, data interfaces and configuration, helping teams turn exploratory work into repeatable components whose dependencies and outputs can be understood and tested.

DVC
Versioning & Lineage

DVC manages versioned data and model artifacts alongside version-controlled project metadata. It links large files stored outside Git to their recorded identities and can describe pipeline dependencies, helping teams reproduce experiments and compare results without committing every dataset byte to the source repository.

Dagster
Workflow Orchestration

Dagster is an orchestration platform that emphasizes data assets and the computations that produce them. The skill defines asset dependencies, partitions and checks so teams can understand what data exists, how it was created and which work is needed to update or repair it.

Prefect
Workflow Orchestration

Prefect is a Python orchestration engine for coordinating flows and tasks with recorded execution state. The skill turns ordinary Python workflows into supervised operations, defining retries, dependencies and deployment behavior while retaining the flexibility to create work dynamically from runtime data and conditions.

Workflow Orchestration
Workflow Orchestration

Workflow orchestration coordinates dependent units of work and supervises their execution over time. It manages scheduling, state, retries and recovery so a data or AI process can complete predictably, while keeping the meaning and correctness of each task in the code and systems that perform it.

Data Cleaning
Data Quality

Data cleaning detects and corrects records that are missing, inconsistent, duplicated, malformed or otherwise unsuitable for an intended use. It combines explicit rules with domain judgment, preserving the difference between a genuine unusual observation and an error that should be repaired or excluded.

PII Redaction
Privacy & PII

PII redaction removes or obscures identifying information before content is shared or processed. The skill combines detection, transformation and residual-risk evaluation, aiming to protect people while retaining enough meaning for the permitted task and recognizing that removing obvious identifiers does not necessarily make a record anonymous.