Atlas · skill

Apache Iceberg

Apache Iceberg is an open table format for large analytical datasets stored as files. It adds table metadata, snapshots and evolution mechanisms above the file layer, allowing compatible engines to read consistent table states without treating a directory listing as the complete definition of a table.

toolTable Formats & Storage

What it is

A file format describes individual files, while a table format describes how files together form a changing table. Iceberg tracks schemas, manifests and snapshots, with commits publishing a new table state. Readers can use a consistent snapshot while writers update the table. Metadata supports features such as schema and partition evolution without requiring users to encode every change in directory conventions. A catalog helps locate and coordinate table metadata, while query engines perform computation. The format does not itself supply all governance or processing capabilities; behavior depends on compatible engines, catalog configuration and the operations they support.

What the work involves

The practitioner chooses a catalog and compatible engines, defines tables and tests read/write behavior across the intended stack. They design file sizes and partitioning from query patterns, then plan maintenance for compaction, snapshots and unused files. Useful artifacts include compatibility tests and a retention policy for historical snapshots. Concurrent writes and failed commits require testing, especially when multiple engines participate. The team verifies that cleanup preserves files still referenced by valid table states and that evolution behaves consistently for every consumer rather than only the writer used in a demonstration.

Illustrative example

An analytics team adds a new field to a transaction table while readers continue querying older snapshots. It later changes partitioning to suit current access patterns without forcing analysts to reason about mixed directory layouts. A maintenance job compacts small files and expires snapshots under a defined policy. Before enabling another engine, the team checks its support for the table features already used and compares query results.

Limits and common mistakes

An open format does not guarantee that every engine supports every feature identically. Small files, poor partitioning and neglected metadata maintenance can still make queries expensive. Snapshot expiration also affects reproducibility and recovery. The practitioner should distinguish table-format guarantees from catalog, storage and query-engine behavior, and test the actual combination. Iceberg is not a replacement for data modeling, access controls or meaningful quality validation.

Prerequisites

  • mediumSQL

    Table formats are queried with SQL — understanding SQL helps leverage their capabilities

Related skills

  • → is an instance of: Data Engineering

Sources and further reading

Last updated: 2026-10-10