Reference · print me

Glossary

The vocabulary this course uses, defined once. Terms marked spec use the Iceberg specification’s own wording.

Source: Iceberg table spec § Terms, plus the official docs.

1.The structure

TermDefinition
Table formatA protocol for agreeing which files constitute a table, layered over immutable files in object storage. It is not a file format — Parquet is the file format underneath.
Schema (spec)Names and types of fields in a table.
Partition spec (spec)A definition of how partition values are derived from data fields.
Snapshot (spec)The state of a table at some point in time, including the set of all data files.
Manifest list (spec)A file that lists manifest files; one per snapshot.
Manifest (spec)A file that lists data or delete files; a subset of a snapshot. Not tied to a partition.
Data file (spec)A file that contains rows of a table.
Delete file (spec)A file that encodes rows of a table that are deleted by position or data values.
Table metadata fileThe JSON document holding schemas, partition specs, sort orders, properties and the snapshot list. A new one is written on every commit.
CatalogThe service that holds the pointer to the current table metadata file and performs the atomic swap that commits it. The only mutable state in the system.
PuffinIceberg’s blob file format, used for table statistics and — since v3 — deletion vectors.

2.Commit and concurrency

TermDefinition
Atomic swapReplacing the catalog’s metadata pointer in one indivisible operation. The basis for serializable isolation.
Optimistic concurrencyEach writer assumes no one else is writing, does the work, then finds out at commit time. Losers retry; nobody locks.
Check-and-putThe catalog operation that validates the base version is still current before swapping. What rejects a stale commit.
Retry validationStructuring a commit as assumptions + actions, so after a conflict the writer re-checks the assumptions and re-applies the actions rather than redoing the work.
Sequence numberA monotonically increasing long assigned to every successful commit. Manifests, data files and delete files inherit their snapshot’s. Determines which deletes apply to which data.
InheritanceWriting null for a sequence number (or row ID) and filling it in from the parent metadata at read time — so a retry rewrites only the manifest list.
Serializable isolationAll table changes occur in a linear history of atomic updates; readers always see a committed snapshot and hold no lock.
Snapshot referenceA named pointer to a snapshot. A tag labels one snapshot; a branch is mutable and advances as you commit to it.
Time travelReading a snapshot that is no longer current. Available exactly as long as retention keeps that snapshot.
Expire snapshotsDropping old snapshots from the table’s history so their unreferenced files can be deleted. Also the operation that ends your ability to travel back to them.

3.Layout and evolution

TermDefinition
Hidden partitioningIceberg derives partition values from a column via a transform and tracks the relationship, so no partition column exists and no query names one.
Partition tupleThe derived partition values for a data file, stored in that file’s manifest entry — beside the file, not inside it.
TransformThe function producing a partition value: identity, bucket[N], truncate[W], year, month, day, hour, void. v3 allows multiple source columns.
Inclusive projectionConverting a scan predicate into a partition predicate such that if the scan predicate matches a row, the partition predicate matches that row’s partition. May over-include; can never under-include.
Scan planningChoosing which files a query must read, by pruning at the manifest list, then the manifest, then the file footer.
Field IDThe unique integer identifying a column. Data files store IDs, not names — which is why rename, drop and reorder are free.
Type promotionWidening a column’s type in place. Narrowing is never allowed.
Partition evolutionChanging the partition spec of an existing table. Metadata-only; old data keeps its old spec.
Split planningPlanning each partition layout separately, with the filter derived for that layout, so several specs coexist in one table.
Sort orderA declared ordering of rows within data files, recorded per file so engines can exploit it.
CompactionRewriting many small data files into fewer larger ones. A normal commit, so it is rollback-able — and it does move bytes.

4.Deletes and v3

TermDefinition
Copy-on-write (COW)Deleting or updating rewrites the affected data files. Slow writes, free reads.
Merge-on-read (MOR)Deleting writes a delete file instead. Fast writes; readers apply the deletes at scan time.
Position deleteA delete identified by data file path and row position. Deprecated in v3 in favour of deletion vectors.
Equality deleteA delete identified by column values, e.g. id = 5. Cheap to write, expensive to apply — and row lineage is not tracked for rows it updates.
Deletion vector (DV)v3’s replacement for position deletes: deleted positions in a Roaring bitmap, stored as a deletion-vector-v1 Puffin blob. At most one per data file per snapshot.
Row lineagev3 tracking of _row_id and _last_updated_sequence_number per row, assigned by inheritance from the snapshot’s first-row-id.
VARIANTA type for semi-structured data whose shape varies per row. Type and binary encoding are defined in the Parquet project, not in Iceberg.
ShreddingStoring frequently-accessed attributes of a VARIANT as physical Parquet columns at write time, so queries avoid parsing.

5.Ecosystem

TermDefinition
LakehouseOpen table formats over object storage, governed by a catalog, readable by many engines — warehouse semantics without a warehouse’s storage lock-in.
Iceberg REST Catalog (IRC)The standard HTTP protocol for talking to an Iceberg catalog. Speaking it is what lets an engine read a table with no bespoke connector.
Catalog CommitsDelta Lake’s adoption of catalog-coordinated commits — the same model Iceberg has always used. A Databricks-led effort; check its status in the Delta project before relying on it.
UniFormDatabricks’ mechanism for exposing a Delta table’s metadata as Iceberg so Iceberg clients can read it. Vendor feature, not part of either spec.
Adaptive Metadata TreeThe proposed unified metadata layer under Delta 5.0 and Iceberg v4. Roadmap, not spec — the ratified v4 change so far is relative locations.
Orphan filesFiles in storage that no snapshot references — usually the residue of failed commits. Cleaning them up is a separate maintenance job from expiring snapshots.

Companion page: Anatomy of an Iceberg table · Course: Open Table Formats → Apache Iceberg