Apache Iceberg
Apache Iceberg is an open table format for large analytical datasets stored as files. It adds table metadata, snapshots and evolution mechanisms above the file layer, allowing compatible engines to read consistent table states without treating a directory listing as the complete definition of a table.
What it is
A file format describes individual files, while a table format describes how files together form a changing table. Iceberg tracks schemas, manifests and snapshots, with commits publishing a new table state. Readers can use a consistent snapshot while writers update the table. Metadata supports features such as schema and partition evolution without requiring users to encode every change in directory conventions. A catalog helps locate and coordinate table metadata, while query engines perform computation. The format does not itself supply all governance or processing capabilities; behavior depends on compatible engines, catalog configuration and the operations they support.
What the work involves
The practitioner chooses a catalog and compatible engines, defines tables and tests read/write behavior across the intended stack. They design file sizes and partitioning from query patterns, then plan maintenance for compaction, snapshots and unused files. Useful artifacts include compatibility tests and a retention policy for historical snapshots. Concurrent writes and failed commits require testing, especially when multiple engines participate. The team verifies that cleanup preserves files still referenced by valid table states and that evolution behaves consistently for every consumer rather than only the writer used in a demonstration.
Illustrative example
An analytics team adds a new field to a transaction table while readers continue querying older snapshots. It later changes partitioning to suit current access patterns without forcing analysts to reason about mixed directory layouts. A maintenance job compacts small files and expires snapshots under a defined policy. Before enabling another engine, the team checks its support for the table features already used and compares query results.
Limits and common mistakes
An open format does not guarantee that every engine supports every feature identically. Small files, poor partitioning and neglected metadata maintenance can still make queries expensive. Snapshot expiration also affects reproducibility and recovery. The practitioner should distinguish table-format guarantees from catalog, storage and query-engine behavior, and test the actual combination. Iceberg is not a replacement for data modeling, access controls or meaningful quality validation.
Prerequisites
- mediumSQL
Table formats are queried with SQL — understanding SQL helps leverage their capabilities
Related skills
- → is an instance of: Data Engineering
Sources and further reading
- Apache Iceberg: introduction
Official open table format, snapshot and schema evolution documentation.
Last updated: 2026-10-10