Lesson 5 · Open table formats
v1 → v4, and the convergence
What each format version added, what v3 changed about deletes, and why “Delta or Iceberg?” is a different question in 2026 than it was in 2022.
1.Four versions, four problems
The format version number goes up only when a change would break forward compatibility — when older readers would not read newer tables correctly.1 So each bump marks something genuinely new on disk.
- v1 answered Lesson 1’s question: make the table an explicitly tracked list of immutable files, with snapshots, hidden partitioning and schema evolution.
- v2 added row-level updates and deletes — encoding which rows are gone without rewriting the files that hold them.
- v3 extended the type system and reworked deletes. It is complete and adopted by the community.
- v4 restructures metadata, starting with relative locations so a table can be moved without rewriting every path. Under active development, not yet adopted.
- Upgrading is one-directional. Before you bump a table to v3, every engine that reads it must understand v3.
2.v2: the two kinds of delete
Deleting a row from an immutable file means recording the deletion somewhere else. v2 gave two ways.1
Position deletes
Mark a row deleted by data file path + row position. Precise, and cheap to apply, because the reader knows exactly which ordinal to skip.
Equality deletes
Mark a row deleted by column values — id = 5. Cheap to write (you needn’t find the row first), expensive to apply, since every candidate file must be checked.
This is the merge-on-read / copy-on-write choice you set per table. Copy-on-write rewrites the affected data files at delete time — slow writes, fast reads. Merge-on-read writes a delete file — fast writes, and readers pay to apply it.
Sequence numbers are what make it safe: a delete applies to a data file only when the data file’s sequence number is less than or equal to the delete’s.1 That is the Lesson 3 mechanism doing its real job.
3.v3: what actually changed
Deletion vectors replace position delete files
A deletion vector encodes deleted positions in a bitmap — a set bit at position P means the row at P is deleted.1 The details are worth knowing because they explain the performance claim:
- Stored as a
deletion-vector-v1blob in a Puffin file.1 - Built from Roaring bitmaps: 64-bit positions split into a 32-bit key and a 32-bit sub-position, with one bitmap per key.1
- Delete manifests track a DV by
file_path,content_offsetandcontent_size_in_bytes, so several DVs can share one Puffin file.1 - At most one DV per data file per snapshot, and writing one must replace all previously written position delete files for that data file.1
That last rule is the design. One delete structure per file means a reader does one lookup instead of merging an unbounded pile of delete files — and it makes comparing the previous and current delete state a straightforward diff. Position delete files are now deprecated: existing ones stay valid, but updating deletes for a data file must produce a DV.1
Row lineage
In v3 and later, tables must track row lineage for all newly created rows.1 Two fields do the work:
| Field | Meaning |
|---|---|
_row_id | A unique long identifying a row within the table. |
_last_updated_sequence_number | The sequence number of the commit that last updated that row. |
Both are assigned by inheritance — from the snapshot’s first-row-id down through manifests — because neither the commit sequence number nor the starting row ID is known until the snapshot actually commits.1 The same trick as sequence numbers in Lesson 3, applied one level finer.
The caveat to remember
Row lineage is not tracked for rows updated via equality deletes.1 If you are counting on lineage for incremental processing or change feeds, that is the case that will surprise you.
And the type system
| Addition | Why it exists |
|---|---|
variant | Semi-structured data whose shape differs per row. Neither primitive nor nested; the encoding is defined in the Parquet project, not in Iceberg.1 |
geometry, geography | Spatial analysis without a bespoke encoding per engine. |
timestamp_ns, timestamptz_ns | Nanosecond precision, for telemetry and financial data. |
unknown | A type that promotes to anything — useful for a column whose type is not yet decided. |
| Default values | Add a column with a value for existing rows, still without rewriting files. |
| Multi-argument transforms | Derive one partition or sort field from several source columns. |
| Table encryption keys | Encryption as table metadata rather than a storage-layer concern. |
Notice where variant is defined. The types that matter most are being standardised in Parquet and then adopted by both table formats — which is exactly the shape of the next section.
4.The convergence
Through 2022 the honest answer to “Delta or Iceberg?” involved real capability gaps. It mostly doesn’t any more, and the reason is that each format adopted the other’s best idea.
- Two formats, two metadata trees. For years, Delta consumers and Iceberg consumers spoke different catalog APIs, and bridging them meant copying and converting data.3
- Catalog as commit coordinator — Iceberg’s model from the start. Delta Lake adopts it with Catalog Commits, which is what unlocks multi-table transactions and consistent governance across engines.2
- Deletion vectors travelled the other way: Delta had them, Iceberg added them in v3.2
- Row lineage, likewise — new to Iceberg in v3, aligning with Delta. Together with DVs it is what makes incremental processing practical.2
- VARIANT and the geospatial types arrived in both from underneath: standardised in Parquet, then adopted by Delta and Iceberg.2
- The stated direction is unifying the metadata layer itself, so one table can be read and written natively by both families of engine.2
Keep two claims apart
The spec says v4 “restructures metadata for improved performance and new capabilities,” and the concrete change it lists is relative locations. It also says v4 is under active development and has not been formally adopted.1
The deck describes a broader “Adaptive Metadata Tree” shared by Delta 5.0 and Iceberg v4, with a Q3 preview.2 That is a vendor roadmap, not a ratified spec. Cite it as direction; never as behaviour you can rely on.
5.So: Delta or Iceberg?
Run the decision on what still differs, not on a feature matrix that has mostly closed.
| Decide on | Because |
|---|---|
| Your catalog | Lesson 3: the catalog owns the only mutable state and the only atomic operation. It constrains who can write, how governance works, and which engines can join. This is the real decision. |
| Which engines must write | Reading is broadly solved. Concurrent writing from several engines is where support is uneven, and where a bad choice hurts for years. |
| Operational maturity in your stack | Compaction, snapshot expiry, orphan-file cleanup. Who runs them, and what happens when nobody does. |
| Format version floor | v3 is only usable if every reader in your estate understands v3. Survey that before you upgrade, not after. |
| Present in both. In 2026 these no longer separate the two formats. |
The line worth being able to say out loud
“In 2026 the format is rarely the constraint — the catalog is. Pick the catalog that fits your governance and your writers, then use whichever table format that catalog supports best. And where you genuinely need both, the Iceberg REST Catalog is the interoperability seam, because it lets an engine read a table without a bespoke connector.”3
6.Where the formats go after v4
Direction, not spec. Four proposals aimed at AI-era workloads, all being worked across the Parquet, Delta and Iceberg communities:2
| Proposal | Problem it targets | Status per the deck |
|---|---|---|
| VARIANT + shredding | Deserialising semi-structured data on the fly is slow at scale; shredding stores common attributes as physical Parquet columns. | Shipped in v3 |
FILE data type | PDFs, images and audio live outside tables today, losing governance, lineage and joinability with structured rows. | Private preview; proposal in Parquet |
| Column append | Adding a column (an embedding, say) rewrites entire Parquet files. Sidecar files joined back at read time instead. | Proposal in Parquet |
| Indexing | Vector and full-text indexes defined on the table, so RAG workloads stop duplicating data into a separate vector store. | Early / in progress |
Note the pattern across all four: the change lands in Parquet first, then both table formats adopt it. If you want a leading indicator of where Delta and Iceberg are going, read the Parquet proposals.
7.Where to go from here
Primary source: the Iceberg table spec, § Format Versioning, § Deletion Vectors, § Row Lineage. Then the Puffin spec for what a deletion vector physically is.
Then stop reading and go argue. Everything in these five lessons is knowledge; the judgement calls — how often to compact, when merge-on-read stops paying, whether to bump to v3 yet — only come from people running it. Two places worth your time:
- The Apache Iceberg Slack — the project’s own workspace, where committers answer. The right venue for “is this the intended behaviour?”
- The
dev@iceberg.apache.orglist — where spec proposals are actually argued. Read the v4 threads before forming an opinion about v4.
And keep the two reference pages: Anatomy of an Iceberg table and the glossary.
Sources
- Apache Iceberg Table Spec — § Format Versioning (v1–v4), § Row-level Deletes, § Deletion Vectors, § Position Delete Files, § Row Lineage, § Scan Planning.
- Benjamin Mathew & Nishith Agarwal, Your Guide to Open Table Formats: Delta, Iceberg, and What’s Next!, Data + AI Summit 2026 — §2 format co-evolution, §2b adaptive metadata tree, §4 formats for AI, §5 roadmap. Vendor deck; GA claims are Databricks GA.
- Tia Chang & Balaji Swamynathan, Apache Iceberg Interoperability: First-Class Support in Databricks OpenSharing, Data + AI Summit 2026 — §the ecosystem split, §native Iceberg reads.