A five-lesson course
Open Table Formats → Apache Iceberg
From “why isn’t a directory of Parquet files a table?” to the 2026 state of play — the metadata tree, the commit protocol, evolution without rewrites, and what v3 and v4 actually change.
1.Who this is for
You are comfortable with Parquet and SQL, and the lakehouse layer above them is new. The course starts one level below Iceberg’s API — at the problem it was built to solve — and never asks you to take a mechanism on trust.
You will be able to
- Draw the metadata tree and say what each layer prunes
- Explain a commit, including when two writers race
- Predict which operations rewrite data and which don’t
- Answer “Delta or Iceberg?” with the 2026 answer
How it is built
- Every sequential idea is an animated scene you can step through
- Prose is kept short on purpose — the diagrams carry the argument
- No quizzes
- Each page prints cleanly, animations and all
2.The lessons
| Lesson | The one thing it gives you | Time | |
|---|---|---|---|
| 1 | A table, not a directory | Exactly what a table format adds over Parquet — and the four silent failures it removes. | 9 min |
| 2 | The metadata tree | The four layers between the catalog and your data files, and how a query collapses through them. | 10 min |
| 3 | The commit, and time travel | Optimistic concurrency, retry validation, sequence numbers — and why the catalog is the real decision. | 10 min |
| 4 | Evolution without rewrites | Hidden partitioning, field IDs, split planning; which changes are free and which move petabytes. | 10 min |
| 5 | v1 → v4, and the convergence | What each version added, what v3 changed about deletes, and how to make the format decision now. | 11 min |
3.Reference
Two pages built to be printed and kept next to you, not read once.
- Anatomy of an Iceberg table — the metadata tree, the field names, the commit protocol, the transforms, the version matrix, and what rewrites data.
- Glossary — every term the course uses, with the spec’s own wording where it has one.
4.A note on sources
Mechanics come from the Iceberg table specification and the official docs — there is no canonical Iceberg academic paper, so the spec is the paper. The 2026 state of play in Lesson 5 additionally draws on two Data + AI Summit 2026 sessions; those are vendor decks, and the lesson says so wherever it uses one.
Where a claim is contested between a vendor deck and the spec, the spec wins and the lesson shows both. The clearest instance is v4: the spec says not formally adopted.