Lesson 2 · Open table formats

The metadata tree

Four layers between the catalog and your Parquet files. Once you can name them and say what each one prunes, Iceberg stops being magic.

~10 min · grounded in the Iceberg spec, § Manifests, § Snapshots, § Scan Planning

2 of 5

1.Build the tree

Lesson 1 ended with a pointer and a snapshot. Zoom in: between “the catalog’s pointer” and “a Parquet file” there are four distinct files, each doing one job.

CATALOG current metadata pointer One pointer. The only thing that changes on a commit. v3.metadata.json table metadata file schemas · partition-specs sort-orders · properties snapshots · current-snapshot-id snap-8273.avro manifest list · one per snapshot per-manifest partition summaries + file counts m-001.avro manifest m-002.avro manifest one row per data or delete file: partition tuple + column metrics a.parq b.parq c.parq d.parq Parquet / Avro / ORC immutable once written A commit rewrites the top of the tree and reuses the bottom. New metadata.json + new manifest list · unchanged manifests are pointed at again, not rewritten
  1. The catalog holds one pointer to the current table metadata file. This is the only mutable thing in the entire structure.
  2. The table metadata file is JSON. It holds the schemas, the partition specs, the sort orders, the properties, and the full list of snapshots.
  3. The manifest list — one file per snapshot — names the manifests in that snapshot, with summary stats so you can skip whole manifests.
  4. A manifest is an Avro file with one row per data or delete file: its path, its partition tuple, and its column metrics.
  5. At the bottom, the actual data files. Immutable. Nothing ever edits one in place.
  6. Because manifests are files, an unchanged one is simply referenced again by the next snapshot. Commits stay cheap even on huge tables.
mutable pointer metadata file per-snapshot
The spec calls this “a persistent tree structure”: every write produces a new snapshot that reuses as much of the previous snapshot’s metadata tree as possible.2

2.What each layer actually holds

The field names matter — you will meet them in metadata tables, in engine logs, and in every debugging session.

LayerKey fieldsIts one job
Table metadata
v3.metadata.json
format-version, table-uuid, location, last-sequence-number, schemas, current-schema-id, partition-specs, sort-orders, snapshots, current-snapshot-id1 Define the table: its shape, its history, and which snapshot is current.
Snapshot
(inside the metadata JSON)
snapshot-id, parent-snapshot-id, sequence-number, timestamp-ms, manifest-list, summary, schema-id, and in v3 first-row-id1 Name one complete, committed state of the table.
Manifest list
snap-*.avro
manifest_path, manifest_length, partition_spec_id, content (0 = data, 1 = deletes), sequence_number, min_sequence_number, added_snapshot_id, added_files_count, partitions1 Let a planner skip entire manifests without opening them.
Manifest
m-*.avro
Per entry: status (added / existing / deleted), snapshot_id, sequence_number, and a data_file struct with the path, partition tuple, record count, and per-column bounds1 Let a planner pick individual files.

Two things worth noticing

Manifests are not partitions. The spec is explicit: manifests “can track data files with any subset of a table and are not associated with partitions.”1 A manifest is a unit of metadata packaging, not of data layout.

Delete files live in the same structure. A manifest’s content field says whether it tracks data files or delete files. Row-level deletes are not a bolt-on — they are first-class tree citizens. Lesson 5 gets into what is inside them.

3.How a query collapses

This is where the layers earn their keep. Take a filter over a large table and watch it prune at every level before a single byte of Parquet is read.

SELECT count(*) FROM logs
WHERE dt = DATE '2026-08-05'
  AND status = 500;
Snapshot every live file 12,400 files Manifest list partition summaries 3 of 40 manifests Manifest rows partition tuple 1,900 files Manifest rows column bounds 9 files Parquet footer row-group stats 23 row groups Four of those five prunes happened in metadata · O(1) RPCs, planned by the client
  1. Start with everything the current snapshot tracks: 12,400 data files.
  2. The manifest list carries partition summaries per manifest, so manifests that cannot contain dt = 2026-08-05 are skipped without being opened.
  3. Inside the surviving manifests, the scan predicate is converted to a partition predicate and matched against each file’s partition tuple.
  4. Still in the manifest: per-column upper and lower bounds on status eliminate files that cannot hold a 500.
  5. Only now does Parquet get involved — the same statistics idea you already know, one level down, skipping row groups.
  6. Four of the five prunes happened in metadata, and none of them required a listing call or a round trip to a metastore.
The counts are illustrative; the order and the mechanism are exactly what the spec describes in § Scan Planning.1

One precise term worth keeping: inclusive projection

Converting a scan predicate into a partition predicate uses an inclusive projection: if a scan predicate matches a row, the partition predicate must match that row’s partition.1

For a table partitioned by day(ts), the query ts > X becomes ts_day >= day(X). It is called inclusive because it may pull in extra rows — a file for that day holds rows on both sides of X — but it can never miss one. Pruning that is allowed to be wrong in only one direction is what makes it safe to trust.

4.Why commits stay cheap

The tree is persistent in the data-structures sense: writing a new version does not copy the old one.

What a commit does not do

  • Rewrite existing manifests
  • Touch existing data files
  • List anything

What it does do

  • Write new data files
  • Write a manifest for them
  • Write a new manifest list, mostly of old entries
  • Write a new metadata JSON, and swap the pointer

The spec notes the consequence for retries: appends “usually create a new manifest file for the appended data files, which can be added to the table without rewriting the manifest on every attempt.”2 Work survives a failed commit. Lesson 3 is about exactly that failure.

The operational cost this shifts, rather than removes

Cheap commits mean many commits, and every one adds a manifest list, at least one manifest, and some data files. Left alone, a streaming table accumulates thousands of small files and a manifest tree that is expensive to read. That is why table maintenance — compaction, manifest rewriting, snapshot expiry — is not optional housekeeping on Iceberg. It is the other half of the design.

5.Read this next

Primary source: the Iceberg table spec, sections Overview, Manifests, Manifest Lists and Scan Planning. Read the field tables rather than skimming them — they are the actual contract, and they are short.

Then keep Anatomy of an Iceberg table next to you; it is this lesson compressed to one printable page.


Sources

  1. Apache Iceberg Table Spec — § Overview, § Table Metadata Fields, § Snapshots, § Manifest Lists, § Manifests, § Scan Planning.
  2. Apache Iceberg docs › Reliability — persistent tree structure, metadata reuse, cost of retries.