Lesson 1 · Open table formats
A table, not a directory
You already know Parquet. So why isn’t a folder full of Parquet files a table? Answering that precisely is the whole reason Iceberg exists.
1.What Parquet gives you, and what it doesn’t
A Parquet file is a self-describing, statistics-carrying unit of data. Schema, row groups, min/max per column — all of it inside one file. That is why engines can skip row groups without reading them.
What Parquet has no opinion about is which files are the table right now. That question is one level up, and it is where everything goes wrong.
Parquet answers
- What columns and types are in this file?
- What are the min/max for column
xin this row group? - How do I decode these bytes?
Nothing answers
- Which files make up the table as of now?
- What did it look like an hour ago?
- Did that overwrite finish?
The Hive answer was: the files under this directory prefix. Combined with a metastore that tracked partitions, that made a table out of a bucket. It also made the table a lie the moment two things happened at once.
2.The read that lied
- A reader treats the prefix as the table: list the directory, read whatever is there. Four files, one clean answer.
- A writer starts an
INSERT OVERWRITE. There is no transaction here — just files appearing and disappearing. - Mid-flight: two old files are gone, one new file is half-written. The directory is a construction site.
- A reader that lists now gets
a,b,e— a mix of two versions. It returns a confident, wrong answer. - Separately, cost: planning means listing every partition directory, so it scales with the table, not with the query.
3.Four failures, named
The Iceberg docs are blunt about what it was built to fix: “Iceberg was designed to solve correctness problems that affect Hive tables running in S3.”2
| Failure | Why it happens |
|---|---|
| No atomic change | State is split between a metastore (partitions) and a filesystem (files). There is no single thing to swap.2 |
| Wrong results from listing | Eventually consistent stores like S3 may return incorrect results when a table’s state is reconstructed by listing files.2 |
| Planning is O(n) | Job planning makes many slow listing calls — O(n) with the number of partitions.2 |
| Silently wrong partitions | Hive must be given partition values and can’t validate them. Writing 2018-12-01 where 20181201 was meant “produces silently incorrect results, not query failures.”3 |
The pattern
Three of the four are silent. Nothing throws. That is what makes this a format problem rather than an operations problem — you cannot monitor your way out of it.
4.What Iceberg does instead
One sentence from the spec carries the whole design: “This table format tracks individual data files in a table instead of directories.”1 Files exist in storage; they only become the table when a commit says so.
- The catalog holds a single pointer to the current metadata file. That file names a snapshot, and the snapshot names every data file. Reader 1 loads it and is now pinned.
- A writer uploads
eandf. They are real objects in storage — and completely invisible to the table, because nothing points at them yet. - The writer builds
v2.metadata.jsondescribing snapshot S2. Still invisible: the catalog pointer has not moved. - The commit is one atomic swap of the pointer, v1 to v2. There is no in-between state for anyone to observe.
- Reader 1 finishes on S1 — unaffected, and holding no lock. Reader 2 starts and gets S2. That is serializable isolation.
The one-line version
An open table format is a protocol for agreeing which files are the table, layered over immutable files in object storage. Everything else — time travel, schema evolution, concurrent writers — falls out of having that agreement written down.
5.The goals, from the spec itself
These are worth reading once slowly. Each one is a rejection of a specific Hive behaviour.1
| Goal | What it rules out |
|---|---|
| Serializable isolation | Readers observing a write in progress. Reads always use a committed snapshot, and readers will not acquire locks. |
| Speed | O(n) planning. Operations use O(1) remote calls to plan a scan, where n would otherwise grow with partitions or files. |
| Scale | A central metastore as a planning bottleneck. Job planning is handled primarily by clients. |
| Evolution | Migrations to add a column. Full schema and partition spec evolution, including safe add, drop, reorder and rename in nested structures. |
| Dependable types | Per-engine type behaviour. |
| Storage separation | Queries that must name partition values. Reads are planned using predicates on data values. |
| Formats | Evolution rules that differ by file format. |
Goals 1, 2 and 6 are the ones this lesson has shown moving. Goal 6 gets Lesson 4 to itself.
6.What this asks of storage — almost nothing
Here is the detail that explains why Iceberg won on object stores. The spec requires only three operations of a filesystem:1
Required
- In-place write — files are never moved or altered once written
- Seekable reads
- Deletes — for files no longer in use
Explicitly not required
- Random-access writes — data and metadata are immutable until deleted
- Rename — except for catalogs that implement the commit by atomic rename
“These requirements are compatible with object stores, like S3.”1 Notice what that leaves unanswered, though: something still has to perform an atomic swap of the pointer. Object stores do not offer compare-and-swap on a key. That job belongs to the catalog — and it is why the catalog turns out to be the most consequential choice you make. Lesson 3 comes back to it.
7.Read this next
Primary source: Iceberg › Reliability. Two pages, and the highest ratio of insight to words in the whole documentation set. It states the Hive problem and the snapshot answer in one pass, then covers optimistic concurrency, cost of retries, and retry validation — all of which Lesson 3 unpacks.
Keep the glossary open from here on; the vocabulary arrives fast in Lesson 2.
Sources
- Apache Iceberg Table Spec — § Overview, § Goals, § File System Operations.
- Apache Iceberg docs › Reliability — the Hive-on-S3 correctness problems and the snapshot answer.
- Apache Iceberg docs › Partitioning — § Problems with Hive partitioning.