Lesson 2 · Open table formats
The metadata tree
Four layers between the catalog and your Parquet files. Once you can name them and say what each one prunes, Iceberg stops being magic.
1.Build the tree
Lesson 1 ended with a pointer and a snapshot. Zoom in: between “the catalog’s pointer” and “a Parquet file” there are four distinct files, each doing one job.
- The catalog holds one pointer to the current table metadata file. This is the only mutable thing in the entire structure.
- The table metadata file is JSON. It holds the schemas, the partition specs, the sort orders, the properties, and the full list of snapshots.
- The manifest list — one file per snapshot — names the manifests in that snapshot, with summary stats so you can skip whole manifests.
- A manifest is an Avro file with one row per data or delete file: its path, its partition tuple, and its column metrics.
- At the bottom, the actual data files. Immutable. Nothing ever edits one in place.
- Because manifests are files, an unchanged one is simply referenced again by the next snapshot. Commits stay cheap even on huge tables.
2.What each layer actually holds
The field names matter — you will meet them in metadata tables, in engine logs, and in every debugging session.
| Layer | Key fields | Its one job |
|---|---|---|
Table metadatav3.metadata.json |
format-version, table-uuid, location, last-sequence-number, schemas, current-schema-id, partition-specs, sort-orders, snapshots, current-snapshot-id1 |
Define the table: its shape, its history, and which snapshot is current. |
| Snapshot (inside the metadata JSON) |
snapshot-id, parent-snapshot-id, sequence-number, timestamp-ms, manifest-list, summary, schema-id, and in v3 first-row-id1 |
Name one complete, committed state of the table. |
Manifest listsnap-*.avro |
manifest_path, manifest_length, partition_spec_id, content (0 = data, 1 = deletes), sequence_number, min_sequence_number, added_snapshot_id, added_files_count, partitions1 |
Let a planner skip entire manifests without opening them. |
Manifestm-*.avro |
Per entry: status (added / existing / deleted), snapshot_id, sequence_number, and a data_file struct with the path, partition tuple, record count, and per-column bounds1 |
Let a planner pick individual files. |
Two things worth noticing
Manifests are not partitions. The spec is explicit: manifests “can track data files with any subset of a table and are not associated with partitions.”1 A manifest is a unit of metadata packaging, not of data layout.
Delete files live in the same structure. A manifest’s content field says whether it tracks data files or delete files. Row-level deletes are not a bolt-on — they are first-class tree citizens. Lesson 5 gets into what is inside them.
3.How a query collapses
This is where the layers earn their keep. Take a filter over a large table and watch it prune at every level before a single byte of Parquet is read.
SELECT count(*) FROM logs
WHERE dt = DATE '2026-08-05'
AND status = 500;
- Start with everything the current snapshot tracks: 12,400 data files.
- The manifest list carries partition summaries per manifest, so manifests that cannot contain
dt = 2026-08-05are skipped without being opened. - Inside the surviving manifests, the scan predicate is converted to a partition predicate and matched against each file’s partition tuple.
- Still in the manifest: per-column upper and lower bounds on
statuseliminate files that cannot hold a 500. - Only now does Parquet get involved — the same statistics idea you already know, one level down, skipping row groups.
- Four of the five prunes happened in metadata, and none of them required a listing call or a round trip to a metastore.
One precise term worth keeping: inclusive projection
Converting a scan predicate into a partition predicate uses an inclusive projection: if a scan predicate matches a row, the partition predicate must match that row’s partition.1
For a table partitioned by day(ts), the query ts > X becomes ts_day >= day(X). It is called inclusive because it may pull in extra rows — a file for that day holds rows on both sides of X — but it can never miss one. Pruning that is allowed to be wrong in only one direction is what makes it safe to trust.
4.Why commits stay cheap
The tree is persistent in the data-structures sense: writing a new version does not copy the old one.
What a commit does not do
- Rewrite existing manifests
- Touch existing data files
- List anything
What it does do
- Write new data files
- Write a manifest for them
- Write a new manifest list, mostly of old entries
- Write a new metadata JSON, and swap the pointer
The spec notes the consequence for retries: appends “usually create a new manifest file for the appended data files, which can be added to the table without rewriting the manifest on every attempt.”2 Work survives a failed commit. Lesson 3 is about exactly that failure.
The operational cost this shifts, rather than removes
Cheap commits mean many commits, and every one adds a manifest list, at least one manifest, and some data files. Left alone, a streaming table accumulates thousands of small files and a manifest tree that is expensive to read. That is why table maintenance — compaction, manifest rewriting, snapshot expiry — is not optional housekeeping on Iceberg. It is the other half of the design.
5.Read this next
Primary source: the Iceberg table spec, sections Overview, Manifests, Manifest Lists and Scan Planning. Read the field tables rather than skimming them — they are the actual contract, and they are short.
Then keep Anatomy of an Iceberg table next to you; it is this lesson compressed to one printable page.
Sources
- Apache Iceberg Table Spec — § Overview, § Table Metadata Fields, § Snapshots, § Manifest Lists, § Manifests, § Scan Planning.
- Apache Iceberg docs › Reliability — persistent tree structure, metadata reuse, cost of retries.