Outcome: Explain Iceberg's metadata tree and use its two killer features: hidden partitioning with in-place evolution, and engine-agnostic catalogs.
1. Snapshots, manifests, manifest lists (0:00–3:00)
Writes create immutable data files; a snapshot points at a manifest list → manifests → files. Readers resolve the current snapshot and read only listed files — atomic commits, serializable isolation, and cheap rollback by pointer move.
2. Hidden partitioning + evolution (3:00–7:00)
Partition values live in metadata, not directory paths, so DAY(ts) → HOUR(ts) rewrites nothing: old and new files coexist under one logical spec, and queries filter on the logical column (ts) instead of path fragments. This is the core Iceberg innovation over Hive-style partitioning.
3. Catalogs and engine choice (7:00–10:00)
The catalog (REST, Hive, Glue) tracks the current snapshot pointer. Because the spec is open and engines read the same metadata, Spark today and Trino/Flink tomorrow share tables. Schema enforcement rejects bad upstream renames at write time (opt-in evolution only).
Key moments
- 2:00 — tracing a read through snapshot → manifests → files
- 5:00 — DAY→HOUR evolution with zero rewrites
- 8:30 — Spark vs Trino reading the same table via REST catalog
Check: Why is changing a Hive partition spec dangerous while changing Iceberg's is safe? Answer in one sentence about paths vs metadata.