Lesson 4 · Open table formats

Evolution without rewrites

Rename a column, repartition a ten-petabyte table, change your mind about the layout — and touch no data files. This is the payoff for everything in Lessons 2 and 3.

~10 min · grounded in the Iceberg spec and the Partitioning / Evolution docs

4 of 5

1.The failure Hive partitioning kept quiet

In Hive, a partition is a column you maintain by hand. The docs list what goes wrong, and the recurring word is silently:2

On write

  • Hive can’t validate partition values — it’s up to the writer to produce the correct one.
  • Wrong format (2018-12-01 for 20181201) → silently incorrect results, not query failures.
  • Wrong source column, or wrong time zone → also incorrect results, also not failures.

On read

  • It’s up to the user to write queries correctly.
  • Users who don’t know the physical layout get needlessly slow queries — Hive can’t translate filters.
  • Working queries are tied to the partitioning scheme, so it can’t be changed without breaking them.

That last one is the trap. The partitioning you chose on day one, before you knew the data, is now load-bearing in every query anyone has written since.

2.Hidden partitioning, both directions

Iceberg’s move is to make partitioning a relationship the table knows about, rather than a column a human maintains: it “produces partition values by taking a column value and optionally transforming it” and “keeps track of the relationship.”2

ON WRITE one incoming row event_time = 10:14:22 level = ERROR PARTITION SPEC day(event_time) identity(level) PARTITION TUPLE (20305, ERROR) stored in the manifest entry No partition column exists. Producers and consumers never see one. ON READ the query a human writes WHERE event_time > :x no partition column named INCLUSIVE PROJECTION day >= day(:x) derived, never typed files pruned In Hive both derivations are the human’s job: supply event_date on write · filter on event_date on read Get either one wrong and you get an answer, not an error.
  1. A row arrives. It has an event_time and a level. It does not have, and will never have, a partition column.
  2. The table’s partition spec says how to derive partition values: day(event_time) and identity(level).
  3. The resulting tuple is stored in the file’s manifest entry — in metadata, beside the file, not inside it.
  4. Now the read side. Someone filters on event_time, the actual column, exactly as they would on any table.
  5. Iceberg derives the partition predicate itself via inclusive projection, and prunes. The user never learned the layout.
  6. The Hive version asks a human to do both derivations correctly, forever, with no validation on either side.
“Queries no longer depend on a table’s physical layout. With a separation between physical and logical, Iceberg tables can evolve partition schemes over time.”2

The transforms you can build a spec from

TransformProducesReach for it when
identitythe value, unmodifiedlow-cardinality categoricals — level, region
year month day houran int offset from the epochtime-ranged queries; pick the granularity your filters use
bucket[N]32-bit Murmur3 hash mod Nhigh-cardinality keys you join or filter on equality — user_id
truncate[W]the value truncated to width Wprefix-ish grouping on strings or numbers
voidalways nullretiring a partition field in a v1 table without reordering the spec

Bucketing is exact and cheap: bucket_N(x) = (murmur3_x86_32_hash(x) & Integer.MAX_VALUE) % N.1 And v3 adds multi-argument transforms, so a partition field can be derived from more than one source column.1

3.Why a rename is free: field IDs

Iceberg tracks every column by a unique ID, not by name and not by position. That one decision is what makes the whole set of schema changes metadata-only.3

SCHEMA v1 1id 2user 3amount names are labels ids are identity DATA FILE written under v1 columns carry field ids 1 · 2 · 3 a rename never touches this file SCHEMA v2 1id 2customer 3amount RENAME user → customer id 2 unchanged The old file resolves by id 2 and reads as customer. No rewrite. DROP amount · id 3 retired 4amount Re-adding the name later gets a fresh id, so id-3 data can never resurface. Track columns by NAME and reusing a name un-deletes the old column’s data. Track them by POSITION and you cannot drop one without renaming every column after it.
  1. Schema v1 has three fields. The data file on the right stores its columns tagged with field IDs, not names.
  2. Rename user to customer. The name in the schema changes; ID 2 does not.
  3. The old file still resolves — by ID — and its values now appear under the new name. Nothing was rewritten.
  4. Drop amount, then add a column called amount later. The new one gets ID 4, so the retired ID-3 data can never reappear.
  5. Both alternatives are broken. Name tracking un-deletes columns; position tracking makes a drop shift every name after it.
The docs state the four guarantees this buys: added columns never read another column’s values; dropping, updating or reordering never changes any other column’s values.3

Type promotion — the only shape change allowed

“Update” means widen, never narrow. The permitted promotions:1

FromToNote
intlong
floatdouble
decimal(P,S)decimal(P',S) where P' > Pwiden precision only — scale is fixed
datetimestamp, timestamp_nsv3+. Promotion to timestamptz is not allowed
unknownany typev3+

A sharp edge worth knowing before you hit it

Iceberg’s Avro manifest format does not store the type of lower and upper bounds, and type promotion does not rewrite existing bounds. So after promoting float to double, older files still carry 4-byte bounds where 8 are now expected.1 Promotion is free, but it is not invisible to the planner.

4.Changing the partitioning of a table that already exists

Because queries never name partition values, the layout can change under them. Old data keeps its old spec; new data uses the new one; both live in the same table.3

ONE TABLE, TWO LAYOUTS spec 0 · month(event_time) everything written through 2025 planned with a monthly filter files untouched by the change spec 1 · day(event_time) everything written from 2026 planned with a daily filter same query, finer pruning SPLIT PLANNING each partition layout plans its own files, with the filter derived for that layout
“Partition evolution is a metadata operation and does not eagerly rewrite files.”3 No sequence here — the two layouts are simultaneous, which is the whole point.

5.The table to actually memorise

This is the practical residue of the lesson: which operations are free, and which move petabytes.

OperationCost
Add / drop / rename / reorder a columnMetadata only. No data files rewritten.3
Widen a type (per the table above)Metadata only — but stale bounds linger in old manifests.
Change the partition specMetadata only. Old data keeps its spec; split planning handles the rest.3
Delete rowsDepends on the table’s mode: copy-on-write rewrites files, merge-on-read writes a delete file. Lesson 5.
Compact / re-sort / rewrite manifestsRewrites data. Deliberate maintenance, not a schema change.
Backfill old data into a new partition layoutRewrites data. Optional — the whole design exists so you don’t have to.

The sentence to take away

Hive coupled the logical table to its physical layout, so every layout decision became permanent. Iceberg breaks that coupling in both directions — writers don’t supply partition values, readers don’t filter on them — and that is what makes evolution a metadata edit instead of a migration project.

6.Read this next

Primary source: Iceberg › Evolution, then Iceberg › Partitioning. Read them in that order — Evolution states the guarantees, Partitioning shows the failure they replaced.

The transforms are on the anatomy cheat sheet if you want them next to you while designing a spec.


Sources

  1. Apache Iceberg Table Spec — § Partitioning, § Partition Transforms, § Bucket Transform Details, § Schema Evolution, § Partition Evolution, § Version 3.
  2. Apache Iceberg docs › Partitioning — § Problems with Hive partitioning, § Iceberg’s hidden partitioning.
  3. Apache Iceberg docs › Evolution — schema evolution operations, correctness guarantees, partition evolution and split planning.