Outcome: Run the streaming half of the mini-lab: broker → topic → Kafka-connector read → checkpointed, watermarked windowed count.
1. Broker and topics (0:00–2:00)
Redpanda is a Kafka-compatible broker (single binary, no ZooKeeper). A topic (shop-clicks, 3 partitions) is an append-only log; consumer offsets track progress. Same API as Kafka, so Spark's Kafka connector works unchanged.
2. Connector options that matter (2:00–4:00)
startingOffsets (earliest vs latest), maxOffsetsPerTrigger (backpressure in dev), failOnDataLoss=false (tolerate retention gaps), and checkpointLocation — delete the checkpoint and the stream replays from startingOffsets, the #1 demo-duplicates cause.
3. The windowed-count query (4:00–6:00)
Parse the JSON value first (from_json + schema — to_timestamp on a raw string column yields NULLs), watermark 10 minutes for lates, group by 5-minute windows, sink to memory/Delta. Kill the stream mid-run, restart, prove identical counts: checkpoint replay done right.
Key moments
- 1:00 — producing the first
shop-clicksevent with rpk - 3:00 — checkpoint delete → duplicate replay demo
- 5:00 — watermark vs window interaction on a late event
Check: Your restarted stream double-counts. Name the two most likely causes (checkpoint, watermark/parse) and how you verify each.