Lesson 4 · Open table formats
Evolution without rewrites
Rename a column, repartition a ten-petabyte table, change your mind about the layout — and touch no data files. This is the payoff for everything in Lessons 2 and 3.
1.The failure Hive partitioning kept quiet
In Hive, a partition is a column you maintain by hand. The docs list what goes wrong, and the recurring word is silently:2
On write
- Hive can’t validate partition values — it’s up to the writer to produce the correct one.
- Wrong format (
2018-12-01for20181201) → silently incorrect results, not query failures. - Wrong source column, or wrong time zone → also incorrect results, also not failures.
On read
- It’s up to the user to write queries correctly.
- Users who don’t know the physical layout get needlessly slow queries — Hive can’t translate filters.
- Working queries are tied to the partitioning scheme, so it can’t be changed without breaking them.
That last one is the trap. The partitioning you chose on day one, before you knew the data, is now load-bearing in every query anyone has written since.
2.Hidden partitioning, both directions
Iceberg’s move is to make partitioning a relationship the table knows about, rather than a column a human maintains: it “produces partition values by taking a column value and optionally transforming it” and “keeps track of the relationship.”2
- A row arrives. It has an
event_timeand alevel. It does not have, and will never have, a partition column. - The table’s partition spec says how to derive partition values:
day(event_time)andidentity(level). - The resulting tuple is stored in the file’s manifest entry — in metadata, beside the file, not inside it.
- Now the read side. Someone filters on
event_time, the actual column, exactly as they would on any table. - Iceberg derives the partition predicate itself via inclusive projection, and prunes. The user never learned the layout.
- The Hive version asks a human to do both derivations correctly, forever, with no validation on either side.
The transforms you can build a spec from
| Transform | Produces | Reach for it when |
|---|---|---|
identity | the value, unmodified | low-cardinality categoricals — level, region |
year month day hour | an int offset from the epoch | time-ranged queries; pick the granularity your filters use |
bucket[N] | 32-bit Murmur3 hash mod N | high-cardinality keys you join or filter on equality — user_id |
truncate[W] | the value truncated to width W | prefix-ish grouping on strings or numbers |
void | always null | retiring a partition field in a v1 table without reordering the spec |
Bucketing is exact and cheap: bucket_N(x) = (murmur3_x86_32_hash(x) & Integer.MAX_VALUE) % N.1 And v3 adds multi-argument transforms, so a partition field can be derived from more than one source column.1
3.Why a rename is free: field IDs
Iceberg tracks every column by a unique ID, not by name and not by position. That one decision is what makes the whole set of schema changes metadata-only.3
- Schema v1 has three fields. The data file on the right stores its columns tagged with field IDs, not names.
- Rename
usertocustomer. The name in the schema changes; ID 2 does not. - The old file still resolves — by ID — and its values now appear under the new name. Nothing was rewritten.
- Drop
amount, then add a column calledamountlater. The new one gets ID 4, so the retired ID-3 data can never reappear. - Both alternatives are broken. Name tracking un-deletes columns; position tracking makes a drop shift every name after it.
Type promotion — the only shape change allowed
“Update” means widen, never narrow. The permitted promotions:1
| From | To | Note |
|---|---|---|
int | long | |
float | double | |
decimal(P,S) | decimal(P',S) where P' > P | widen precision only — scale is fixed |
date | timestamp, timestamp_ns | v3+. Promotion to timestamptz is not allowed |
unknown | any type | v3+ |
A sharp edge worth knowing before you hit it
Iceberg’s Avro manifest format does not store the type of lower and upper bounds, and type promotion does not rewrite existing bounds. So after promoting float to double, older files still carry 4-byte bounds where 8 are now expected.1 Promotion is free, but it is not invisible to the planner.
4.Changing the partitioning of a table that already exists
Because queries never name partition values, the layout can change under them. Old data keeps its old spec; new data uses the new one; both live in the same table.3
5.The table to actually memorise
This is the practical residue of the lesson: which operations are free, and which move petabytes.
| Operation | Cost |
|---|---|
| Add / drop / rename / reorder a column | Metadata only. No data files rewritten.3 |
| Widen a type (per the table above) | Metadata only — but stale bounds linger in old manifests. |
| Change the partition spec | Metadata only. Old data keeps its spec; split planning handles the rest.3 |
| Delete rows | Depends on the table’s mode: copy-on-write rewrites files, merge-on-read writes a delete file. Lesson 5. |
| Compact / re-sort / rewrite manifests | Rewrites data. Deliberate maintenance, not a schema change. |
| Backfill old data into a new partition layout | Rewrites data. Optional — the whole design exists so you don’t have to. |
The sentence to take away
Hive coupled the logical table to its physical layout, so every layout decision became permanent. Iceberg breaks that coupling in both directions — writers don’t supply partition values, readers don’t filter on them — and that is what makes evolution a metadata edit instead of a migration project.
6.Read this next
Primary source: Iceberg › Evolution, then Iceberg › Partitioning. Read them in that order — Evolution states the guarantees, Partitioning shows the failure they replaced.
The transforms are on the anatomy cheat sheet if you want them next to you while designing a spec.
Sources
- Apache Iceberg Table Spec — § Partitioning, § Partition Transforms, § Bucket Transform Details, § Schema Evolution, § Partition Evolution, § Version 3.
- Apache Iceberg docs › Partitioning — § Problems with Hive partitioning, § Iceberg’s hidden partitioning.
- Apache Iceberg docs › Evolution — schema evolution operations, correctness guarantees, partition evolution and split planning.