Outcome: Lay out Parquet files so queries prune instead of scan: columnar encoding, partition keys chosen by filter, and no small-files trap.
1. Why columnar (0:00–2:30)
Parquet stores columns together: dictionary + run-length encoding shrink repeated values, and predicate pushdown skips whole row groups without reading them. An analytics scan touching 2 of 40 columns reads ~5% of the bytes a CSV scan would.
2. Partition by what you filter (2:30–5:30)
Partitioning maps a column to directories (order_date=2026-09-20/), so a date filter skips entire directories. Rules: partition on the column queries filter most (usually date); keep cardinality modest — order_id partitioning is a file-per-value explosion; truncate timestamps to date/hour first.
3. The small-files trap (5:30–8:00)
Thousands of tiny files = driver overhead and slow listings. Causes: default 200 shuffle partitions on small data, streaming micro-batches. Fixes: coalesce/repartition before writes, ≤8 files per date partition, scheduled OPTIMIZE on lakehouse tables.
Key moments
- 1:30 — pushdown skipping row groups
- 4:00 — order_date vs order_id partitioning decision
- 6:30 — diagnosing a 2,000-tiny-files write
Check: A 2 GB table with 5M distinct millisecond timestamps — what do you partition by, and why not the raw timestamp?