Atlas · skill

Data Ingestion

Data ingestion brings records from source systems into a destination where they can be processed or consumed. It defines how delivery, schema changes and progress are tracked, so collecting data includes reliable replay and correction behavior rather than merely moving bytes between services.

conceptData Ingestion

What it is

Ingestion can use periodic snapshots, incremental queries, change-data capture or event streams. Each approach exposes different information about updates and deletions and depends on the source's available interface. A checkpoint records progress, while stable record identifiers or offsets help distinguish new delivery from replay. Raw landing storage may preserve source evidence for later transformation, but it also creates retention and access responsibilities. Ingestion differs from downstream modeling because its primary concern is faithful and dependable transfer. A successful request does not establish complete coverage if pagination, source consistency or interrupted delivery was not handled correctly.

What the work involves

The practitioner defines source boundaries, record identity and a delivery strategy that meets freshness requirements. They implement checkpoints, schema validation and a safe response to malformed or unexpectedly changed records. Useful artifacts include an ingestion contract and tests for retries, interrupted batches and deletion propagation. The team reconciles source and destination coverage where possible and monitors lag. If replay is supported, it verifies that downstream writes are idempotent or deduplicated and records which source state each batch represents rather than silently mixing inconsistent snapshots.

Illustrative example

A pipeline imports customer updates from a paginated API. A timeout after page three initially causes duplicate records on retry. The engineer introduces a stable extraction boundary, checkpointing and upserts keyed by source identity. Tests include a source deletion and a schema addition, making their handling explicit. A reconciliation report then identifies records that were skipped or failed instead of treating the final HTTP success as evidence of complete ingestion.

Limits and common mistakes

Some sources cannot provide a consistent snapshot or reliable change history, limiting what ingestion can guarantee. Aggressive polling can overload the source, while long intervals may miss transient states. Schema inference can silently reinterpret values after drift. Strong ingestion makes these limitations visible and supports recovery with evidence. It does not establish that source records are accurate or suitable for model training merely because they were transferred faithfully.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10