Back to M1 — Stats / Experimentation Rounds

Stats / Experimentation Rounds

Outcome: Solve CUPED + Simpson + permutation items Curated video (Jay Feng): A/B Testing Interview with a Google Data Scientist — https://www.youtube.com/watch?v=2sWVLMVQsu0 (verified live via yt-dlp 2026-09-24). Pointer: 77-derived items (L343/q901, L345/q902, L349/q903); shell: courses/video-scripts/ds-interview-prep/02.md.

13 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. CUPED Mechanics — The instructor introduces variance reduction mathematics and demonstrates computing theta using pre-experiment baseline covariates.
  2. Simpson's Paradox Pitfalls — The instructor illustrates how subgroup mix-shifts cause aggregate experiment metrics to contradict segmented treatment lifts.
  3. Permutation Testing for Skew — The instructor constructs an exact empirical null distribution by shuffling treatment labels to evaluate heavy-tailed metric distributions.
  4. Difference-in-Differences Framework — The instructor derives the two-by-two Difference-in-Differences estimator and discusses validating the parallel trends assumption.
PDF notes

Frequently asked questions

How does CUPED reduce required experiment sample size?

Because sample size is proportional to metric variance, reducing variance by $(1 - \rho^2)$ directly reduces the sample size and test duration needed for target power.

What causes Simpson's Paradox in online A/B testing?

It occurs when unequal traffic allocation ratios interact with heterogeneous conversion rates across confounding segments like mobile versus desktop users.

When is a permutation test preferred over the Mann-Whitney U test?

Permutation tests directly evaluate differences in means on the original scale, whereas Mann-Whitney evaluates rank differences that do not test expected value differences.

What happens if parallel trends fail in a Difference-in-Differences design?

The interaction coefficient will conflate existing baseline trend divergence with the true treatment effect, resulting in a biased and invalid causal estimate.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.