Reference · print me
Glossary
The vocabulary this course uses, defined once. Terms marked spec use the Iceberg specification’s own wording.
1.The structure
| Term | Definition |
|---|---|
| Table format | A protocol for agreeing which files constitute a table, layered over immutable files in object storage. It is not a file format — Parquet is the file format underneath. |
| Schema (spec) | Names and types of fields in a table. |
| Partition spec (spec) | A definition of how partition values are derived from data fields. |
| Snapshot (spec) | The state of a table at some point in time, including the set of all data files. |
| Manifest list (spec) | A file that lists manifest files; one per snapshot. |
| Manifest (spec) | A file that lists data or delete files; a subset of a snapshot. Not tied to a partition. |
| Data file (spec) | A file that contains rows of a table. |
| Delete file (spec) | A file that encodes rows of a table that are deleted by position or data values. |
| Table metadata file | The JSON document holding schemas, partition specs, sort orders, properties and the snapshot list. A new one is written on every commit. |
| Catalog | The service that holds the pointer to the current table metadata file and performs the atomic swap that commits it. The only mutable state in the system. |
| Puffin | Iceberg’s blob file format, used for table statistics and — since v3 — deletion vectors. |
2.Commit and concurrency
| Term | Definition |
|---|---|
| Atomic swap | Replacing the catalog’s metadata pointer in one indivisible operation. The basis for serializable isolation. |
| Optimistic concurrency | Each writer assumes no one else is writing, does the work, then finds out at commit time. Losers retry; nobody locks. |
| Check-and-put | The catalog operation that validates the base version is still current before swapping. What rejects a stale commit. |
| Retry validation | Structuring a commit as assumptions + actions, so after a conflict the writer re-checks the assumptions and re-applies the actions rather than redoing the work. |
| Sequence number | A monotonically increasing long assigned to every successful commit. Manifests, data files and delete files inherit their snapshot’s. Determines which deletes apply to which data. |
| Inheritance | Writing null for a sequence number (or row ID) and filling it in from the parent metadata at read time — so a retry rewrites only the manifest list. |
| Serializable isolation | All table changes occur in a linear history of atomic updates; readers always see a committed snapshot and hold no lock. |
| Snapshot reference | A named pointer to a snapshot. A tag labels one snapshot; a branch is mutable and advances as you commit to it. |
| Time travel | Reading a snapshot that is no longer current. Available exactly as long as retention keeps that snapshot. |
| Expire snapshots | Dropping old snapshots from the table’s history so their unreferenced files can be deleted. Also the operation that ends your ability to travel back to them. |
3.Layout and evolution
| Term | Definition |
|---|---|
| Hidden partitioning | Iceberg derives partition values from a column via a transform and tracks the relationship, so no partition column exists and no query names one. |
| Partition tuple | The derived partition values for a data file, stored in that file’s manifest entry — beside the file, not inside it. |
| Transform | The function producing a partition value: identity, bucket[N], truncate[W], year, month, day, hour, void. v3 allows multiple source columns. |
| Inclusive projection | Converting a scan predicate into a partition predicate such that if the scan predicate matches a row, the partition predicate matches that row’s partition. May over-include; can never under-include. |
| Scan planning | Choosing which files a query must read, by pruning at the manifest list, then the manifest, then the file footer. |
| Field ID | The unique integer identifying a column. Data files store IDs, not names — which is why rename, drop and reorder are free. |
| Type promotion | Widening a column’s type in place. Narrowing is never allowed. |
| Partition evolution | Changing the partition spec of an existing table. Metadata-only; old data keeps its old spec. |
| Split planning | Planning each partition layout separately, with the filter derived for that layout, so several specs coexist in one table. |
| Sort order | A declared ordering of rows within data files, recorded per file so engines can exploit it. |
| Compaction | Rewriting many small data files into fewer larger ones. A normal commit, so it is rollback-able — and it does move bytes. |
4.Deletes and v3
| Term | Definition |
|---|---|
| Copy-on-write (COW) | Deleting or updating rewrites the affected data files. Slow writes, free reads. |
| Merge-on-read (MOR) | Deleting writes a delete file instead. Fast writes; readers apply the deletes at scan time. |
| Position delete | A delete identified by data file path and row position. Deprecated in v3 in favour of deletion vectors. |
| Equality delete | A delete identified by column values, e.g. id = 5. Cheap to write, expensive to apply — and row lineage is not tracked for rows it updates. |
| Deletion vector (DV) | v3’s replacement for position deletes: deleted positions in a Roaring bitmap, stored as a deletion-vector-v1 Puffin blob. At most one per data file per snapshot. |
| Row lineage | v3 tracking of _row_id and _last_updated_sequence_number per row, assigned by inheritance from the snapshot’s first-row-id. |
| VARIANT | A type for semi-structured data whose shape varies per row. Type and binary encoding are defined in the Parquet project, not in Iceberg. |
| Shredding | Storing frequently-accessed attributes of a VARIANT as physical Parquet columns at write time, so queries avoid parsing. |
5.Ecosystem
| Term | Definition |
|---|---|
| Lakehouse | Open table formats over object storage, governed by a catalog, readable by many engines — warehouse semantics without a warehouse’s storage lock-in. |
| Iceberg REST Catalog (IRC) | The standard HTTP protocol for talking to an Iceberg catalog. Speaking it is what lets an engine read a table with no bespoke connector. |
| Catalog Commits | Delta Lake’s adoption of catalog-coordinated commits — the same model Iceberg has always used. A Databricks-led effort; check its status in the Delta project before relying on it. |
| UniForm | Databricks’ mechanism for exposing a Delta table’s metadata as Iceberg so Iceberg clients can read it. Vendor feature, not part of either spec. |
| Adaptive Metadata Tree | The proposed unified metadata layer under Delta 5.0 and Iceberg v4. Roadmap, not spec — the ratified v4 change so far is relative locations. |
| Orphan files | Files in storage that no snapshot references — usually the residue of failed commits. Cleaning them up is a separate maintenance job from expiring snapshots. |