Tests are pass/fail rules. Observability is broader: continuous visibility into how data flows through the stack and what's "normal".
What data observability covers
1. Lineage
Where did this data come from? What downstream consumes it?
When customers.lifetime_value changes definition, which dashboards break?
Tools: dbt's docs (column-level lineage), OpenLineage (cross-tool lineage), commercial platforms.
2. Freshness
When was each table last updated? Is the freshness within SLA?
marts.daily_orders last updated: 2026-05-25 03:47
expected: within 24h ✓
marts.daily_active_users last updated: 2026-05-23 03:47
expected: within 24h ✗ ALERT
3. Volume
How many rows arrived today? Is this within historical range?
raw.events_2026_05_25: 1.2M rows (avg of past 7 days: 1.5M, std: 200K)
→ -15% from average, within 1 std → OK
raw.events_2026_05_24: 100K rows (avg: 1.5M, std: 200K)
→ -93% from average, -7 std → ALERT
4. Schema drift
Did columns get added/removed? Types change?
raw.customers added column: 'loyalty_tier' → ALERT (review downstream)
raw.customers removed column: 'email' → CRITICAL (downstream broken)
5. Distribution drift
Are value distributions changing?
raw.orders.country distribution:
Today: US 60%, IN 30%, UK 10%
Last week avg: US 45%, IN 40%, UK 15%
→ US share up significantly; investigate
The tools landscape
OpenLineage
Open standard for lineage metadata. dbt, Airflow, Spark all emit OpenLineage events.
Use case: build your own lineage graph from emitted events.
Free, open, requires plumbing.
Monte Carlo
Commercial data observability platform. Auto-detects freshness, volume, schema, distribution issues across your warehouse.
Use case: get observability without building it yourself.
Pricing: enterprise.
Datafold
Diff-based observability. When dbt models change, Datafold shows the data diff (rows changed, columns differ, distribution shifts).
Use case: CI for dbt — see data impact of PR before merging.
Soda
SQL-first data quality + monitoring. Define quality rules in YAML; runs continuously.
- name: orders_freshness
checks:
- max(ingestion_ts) > now() - 1h
Use case: lighter weight than Monte Carlo; SQL-defined.
dbt-only stack
dbt docs + dbt tests + Slack alerts. Free.
Use case: small teams; covers basics.
What to monitor (the 80/20)
For every important production table:
- Freshness — alert if not updated within expected interval.
- Row count — alert if today's rows are <X% of historical average.
- Null rates on critical columns — alert if >X% nulls in customer_id, etc.
- Distinct value count on PK — alert if uniqueness violated.
These 4 catch ~80% of pipeline issues in practice.
Building it DIY (small budget)
Without commercial tools, you can build basics:
-- monitoring/freshness.sql
CREATE OR REPLACE TABLE monitoring.freshness AS
SELECT
table_name,
MAX(_loaded_at) AS last_load,
DATEDIFF('minute', MAX(_loaded_at), NOW()) AS minutes_since
FROM information_schema.tables
-- ... (join + group by per table)
-- monitoring/alerts.sql (runs hourly)
SELECT * FROM monitoring.freshness
WHERE minutes_since > expected_interval
Send results to Slack via webhook. Crude but works. Many teams start here and add commercial tools later.
Observability vs testing — when to use which
| Aspect | Tests | Observability |
|---|---|---|
| What | Specific rules | Continuous patterns |
| When run | Each pipeline run | Continuous |
| Outcome | Pass/fail | Anomaly score |
| Use | "Email must not be null" | "Email null rate jumped today" |
Use both. Tests are deterministic checks. Observability catches what tests miss.
Lineage in practice
Column-level lineage:
raw.orders.order_total
→ staging.stg_orders.order_total
→ marts.fct_revenue.daily_revenue
→ Looker dashboard "Revenue trend"
→ Email report to CEO
When something changes upstream, you know exactly which downstream artifacts to update.
dbt's docs has column lineage. Commercial tools (Atlan, DataHub) span more sources.
Alerting fatigue
Common failure: 50 alerts/day, no one investigates any.
Patterns to avoid alert fatigue:
- Severity levels. Critical = page; warn = digest.
- Aggregate similar alerts. "5 tables have null spikes" not 5 alerts.
- Owners per table. Alert goes to the owner, not a generic channel.
- Tune thresholds. False alarms = teams turn off alerts.
Common observability mistakes
- No observability. Find out about pipeline issues from dashboards days later.
- Too many alerts. Erodes trust; teams ignore them.
- No lineage. Schema change breaks unknown downstream consumers.
- Spend on Monte Carlo without basics covered. dbt tests + Slack first; commercial second.
- Lineage without ownership. Alerts go to a channel no one reads.
Takeaway
Observability = continuous visibility (freshness, volume, schema, distribution, lineage). Complements specific tests. Tools range from DIY (dbt + custom SQL + Slack) to commercial (Monte Carlo, Datafold, Soda). Start small: monitor freshness + row count + null rates on critical tables. Add lineage and anomaly detection as you grow.