Micro Free Course — 100% Free Learning

All lessons in this module are free to learn. Sign in with Google to save your progress.

2

Reliability and Production Patterns

Data quality testing (Great Expectations, dbt tests); pipeline observability (Monte Carlo, OpenLineage, Datadog); streaming basics (Kafka, Pub/Sub, Kinesis); Spark for batch and structured streaming.

Module Progress0% Complete
222 min total
10 Lessons
0 Completed

Module Content

Data Quality Testing with dbt — Schema Tests to CI/CD

Two dominant approaches to encoding data quality rules. Most production stacks use both. FIND_VIDEO: search 'great expectations vs dbt tests data quality' — recommended channel: Great Expectations / dbt Labs. Aim for 10 min or under.

16 minVideo
Start

Quiz: Data Quality Discipline & Testing

Data quality is the bridge between 'pipeline ran' and 'data is correct'. Encode rules, run them every load, alert on failures.

7 minTutorial
Start

Data Observability — Eliminating Data Downtime (Barr Moses)

Beyond pass/fail tests: end-to-end visibility into pipeline health, lineage, and data freshness. FIND_VIDEO: search 'data observability openlineage monte carlo pipeline' — recommended channel: Monte Carlo / OpenLineage. Aim for 10 min or under.

27 minVideo
Start

Quiz: Data Observability & Incident Response

What 'observability' means in DE: lineage, freshness, volume, schema drift, distribution drift. Tools that surface them.

7 minTutorial
Start

Case 4 — Incident Response: Data Quality SLA Failure Investigation

Investigate a high-severity production data quality incident where revenue dropped 30%, diagnose the root cause, implement tiered dbt tests, and configure freshness alerts.

62 minSubmission
Start

Apache Kafka Architecture — Topics, Partitions & Offsets

The streaming layer of the modern stack. Covers Kafka concepts, when to use streaming, and operational realities. FIND_VIDEO: search 'kafka streaming basics tutorial data engineering' — recommended channel: Confluent / Stephane Maarek. Aim for 10 min or under.

11 minVideo
Start

Quiz: Streaming Fundamentals & Kafka Architecture

Streaming concepts (topics, partitions, consumer groups, offsets) and when streaming is justified vs batch.

7 minTutorial
Start

Case 5 — Architecture Evolution: Batch to Real-Time Streaming Migration

Architect and implement a zero-downtime migration moving an hourly batch fraud detection pipeline to real-time event streaming using Apache Kafka and Python.

62 minSubmission
Start

Apache Spark Fundamentals — Architecture & In-Memory RDDs

When and how to use Spark for batch transformations. The default for large-data ELT outside the warehouse. FIND_VIDEO: search 'apache spark tutorial data engineering' — recommended channel: Apache Spark / Databricks. Aim for 11 min or under.

16 minVideo
Start

Quiz: Distributed Compute — When Spark Belongs in Your Stack

Spark for DE: when warehouse SQL isn't enough; PySpark patterns; Databricks vs open-source Spark; when not to use Spark.

7 minTutorial
Start