Outcome: Place Spark correctly in a modern data stack: what runs in the warehouse, what needs Spark, and how the lakehouse layers (bronze/silver/gold) frame both.
1. The 2-minute Spark mental model (0:00–2:00)
One driver plans; many executors run tasks on partitions. Your DataFrame code is a lazy plan until an action fires it. Everything slow about Spark is a shuffle you didn't plan — narrow transforms (filter, select) stay local, wide ones (groupBy, join) move data.
2. Bronze / silver / gold (2:00–4:30)
- Bronze: raw landed data, faithful copy of source. Rebuildable, never queried by dashboards.
- Silver: cleaned, typed, deduplicated — one row per business event.
- Gold: modeled marts (facts/dims) that dashboards actually query.
The layers are about blast radius: a bug in gold rebuilds from silver, never from the source system.
3. Where Spark fits (4:30–7:00)
Warehouse SQL wins for terabyte-scale declarative analytics. Reach for Spark when: data volume or custom logic (ML features, graph, complex parsing) outgrows warehouse SQL economics; file-format work (Parquet/Delta/Iceberg) outside any warehouse; or streaming (Structured Streaming) with watermarks and checkpoints.
Key moments
- 1:10 — driver vs executors, partitions as parallelism
- 3:20 — why gold never reads bronze directly
- 5:40 — the three "reach for Spark" triggers
Check: Name one workload for warehouse SQL and one for Spark, and say which lakehouse layer each reads from.