This lesson on Stats / Experimentation Rounds is hands-on and example-driven. You will master the core quantitative frameworks tested in data science experimentation rounds: variance reduction via CUPED, diagnosing Simpson's paradox across mix-shifts, exact permutation testing for skewed distributions, and causal estimation using Difference-in-Differences.
What You'll Be Able To Do
- Compute CUPED-adjusted metric values using pre-experiment covariates to accelerate test runtime.
- Diagnose and resolve Simpson's paradox caused by uneven subgroup traffic allocation.
- Implement two-sample permutation tests to evaluate treatment effects on highly skewed metric distributions.
- Evaluate parallel trends assumptions and formulate regression specifications for Difference-in-Differences designs.
Detailed Concept Walkthrough
1. CUPED Variance Reduction
CUPED (Controlled-experiment Using Pre-Experiment Data) removes pre-existing variance in the primary metric by subtracting a linearly scaled pre-experiment covariate. This reduces sample variance without biasing the treatment effect estimator, requiring fewer samples to achieve target statistical power.
- Mechanism: Calculate the optimal adjustment coefficient $\theta = \text{Cov}(Y, X) / \text{Var}(X)$, where $Y$ is the experiment metric and $X$ is the metric measured for the same entity prior to experiment launch. The transformed metric is defined as $\tilde{Y} = Y - \theta(X - \mathbb{E}[X])$.
- Under the Hood: By regressing the post-experiment outcome on the baseline covariate, the variance of $\tilde{Y}$ scales down by a factor of $(1 - \rho^2)$, where $\rho$ is the Pearson correlation between $X$ and $Y$. Higher correlation between baseline and post-treatment metrics yields greater variance reduction.
- Best Practice: Ensure the covariate $X$ is strictly measured before assignment so that it cannot be influenced by the treatment. Using post-treatment covariates introduces post-treatment bias and invalidates causal inference.
import numpy as np
def apply_cuped(y_metric: np.ndarray, x_baseline: np.ndarray) -> np.ndarray:
"""Adjust experiment metric Y using pre-experiment covariate X."""
covariance = np.cov(y_metric, x_baseline)[0, 1]
variance_x = np.var(x_baseline, ddof=1)
theta = covariance / variance_x
y_adjusted = y_metric - theta * (x_baseline - np.mean(x_baseline))
return y_adjusted
Key Takeaway: CUPED scales variance down by $(1 - \rho^2)$ using pre-experiment covariates, directly shrinking required sample sizes and experimentation runtimes.
2. Simpson's Paradox and Mix-Shifts
Simpson's paradox occurs when a trend or effect observed in combined aggregate data reverses or disappears when evaluated across granular sub-populations. It is typically driven by an unobserved confounding variable or unequal traffic split allocations across heterogeneous cohorts.
- Mechanism: When different subgroups have substantially different baseline conversion rates and different treatment-to-control allocation ratios, aggregate averages become weighted sums that reflect cohort composition shifts rather than true treatment effects.
- Under the Hood: If Cohort A (high converting) receives 80% control and 20% treatment, while Cohort B (low converting) receives 20% control and 80% treatment, the overall treatment mean is pulled downward by Cohort B despite positive treatment gains within each individual cohort.
- Best Practice: Always segment experimental metrics across key stratifications (e.g., country, device, user tenure) and run stratified or regression-adjusted models to prevent mix-shift artifacts from misguiding product launch decisions.
import pandas as pd
def check_simpsons_paradox(df: pd.DataFrame, group_col: str, segment_col: str, metric_col: str) -> None:
# Segmented metric
segmented = df.groupby([segment_col, group_col])[metric_col].mean().unstack()
print("Segmented Means:\n", segmented)
# Aggregate metric
aggregate = df.groupby(group_col)[metric_col].mean()
print("Aggregate Means:\n", aggregate)
Key Takeaway: Aggregate experiment results can invert true underlying effects whenever assignment proportions shift across heterogeneous sub-populations.
3. Permutation Testing for Skewed Metrics
Permutation (randomization) tests compute an exact non-parametric empirical null distribution by repeatedly shuffling group assignment labels. This evaluates treatment efficacy without relying on standard central limit theorem normality assumptions.
- Mechanism: Calculate the observed test statistic (e.g., difference in means or medians) between treatment and control. Pool all observations together, randomly permute the assignment labels, and recompute the test statistic across $B$ resampling iterations.
- Under the Hood: The two-sided p-value is computed directly as the fraction of permuted statistics whose absolute value meets or exceeds the observed statistic: $p = \frac{1}{B} \sum_{b=1}^{B} \mathbb{I}(|T_b| \ge |T_{\text{obs}}|)$.
- Best Practice: Use permutation testing when analyzing heavy-tailed or zero-inflated distributions (such as revenue per user or engagement seconds) where standard asymptotic t-tests exhibit inflated Type I or Type II error rates.
import numpy as np
def permutation_test(ctrl: np.ndarray, trt: np.ndarray, n_resamples: int = 10000) -> float:
obs_diff = np.mean(trt) - np.mean(ctrl)
pooled = np.concatenate([ctrl, trt])
n_ctrl = len(ctrl)
# Generate null distribution by permuting labels
diffs = np.empty(n_resamples)
for i in range(n_resamples):
shuffled = np.random.permutation(pooled)
diffs[i] = np.mean(shuffled[n_ctrl:]) - np.mean(shuffled[:n_ctrl])
return float(np.mean(np.abs(diffs) >= np.abs(obs_diff)))
Key Takeaway: Permutation tests construct an exact distribution-free null benchmark by shuffling assignment labels across skewed or zero-inflated metrics.
4. Difference-in-Differences (DiD)
Difference-in-Differences estimates causal treatment effects in non-randomized or quasi-experimental settings by comparing changes over time between a treated group and an untreated control group.
- Mechanism: DiD subtracts the pre-treatment baseline difference from the post-treatment outcome difference: $\hat{\delta} = (\bar{Y}{T,\text{post}} - \bar{Y}{T,\text{pre}}) - (\bar{Y}{C,\text{post}} - \bar{Y}{C,\text{pre}})$.
- Under the Hood: The fundamental identifying assumption is Parallel Trends: in the absence of treatment, the average outcome for the treated group would have followed the same trajectory as the control group. Violations of parallel trends bias the estimated treatment effect $\hat{\delta}$.
- Best Practice: In an interview setting, formulate the DiD estimator via ordinary least squares regression: $Y_{it} = \beta_0 + \beta_1 \text{Treat}_i + \beta_2 \text{Post}_t + \delta (\text{Treat}_i \times \text{Post}t) + \varepsilon{it}$, where $\delta$ captures the causal interaction effect.
import statsmodels.formula.api as smf
import pandas as pd
def fit_did_model(df: pd.DataFrame) -> None:
"""Estimate DiD interaction term via OLS regression."""
# Formula: Y ~ treated + post + treated:post
model = smf.ols('metric ~ treated * post', data=df).fit()
print(model.summary().tables[1])
Key Takeaway: DiD controls for unobserved time-invariant confounders by evaluating divergence in pre-to-post trajectories under the parallel trends assumption.
Topics Covered in Stats / Experimentation Rounds
- CUPED Mechanics (0:00 - 2:30) — The instructor introduces variance reduction mathematics and demonstrates computing theta using pre-experiment baseline covariates.
- Simpson's Paradox Pitfalls (2:31 - 5:15) — The instructor illustrates how subgroup mix-shifts cause aggregate experiment metrics to contradict segmented treatment lifts.
- Permutation Testing for Skew (5:16 - 8:00) — The instructor constructs an exact empirical null distribution by shuffling treatment labels to evaluate heavy-tailed metric distributions.
- Difference-in-Differences Framework (8:01 - 11:00) — The instructor derives the two-by-two Difference-in-Differences estimator and discusses validating the parallel trends assumption.
DS Interview Prep Cheat Sheet
-
CUPED theta parameter— Computes optimal variance-reduction regression coefficienttheta = np.cov(y, x)[0, 1] / np.var(x, ddof=1) -
CUPED transformed metric— Calculates adjusted metric subtracting covariate contributiony_adj = y - theta * (x - np.mean(x)) -
Variance reduction ratio— Quantifies variance scaling as 1 minus squared correlationvar_reduction = 1 - np.corrcoef(y, x)[0, 1]**2 -
Permutation null resample— Draws randomized metric differences under null hypothesisshuffled = np.random.permutation(pooled_data) -
Permutation p-value— Calculates proportion of null draws exceeding observed effectpval = np.mean(np.abs(null_dist) >= np.abs(obs_stat)) -
Difference-in-Differences formula— Estimates treatment interaction parameter in regressionsmf.ols('y ~ treated * post', data=df).fit()
Comparison Table
| Methodology | Primary Use Case | Core Assumption / Requirement |
|---|---|---|
| CUPED | Variance reduction in standard A/B testing | Covariate measured strictly pre-treatment |
| Permutation Test | Hypothesis testing on non-normal distributions | Exchangeability under the null hypothesis |
| Difference-in-Differences | Quasi-experiments and policy evaluation | Parallel trends in counterfactual baseline |
Common Pitfalls
- Mistake: Calculating CUPED adjustment using covariates measured after experiment exposure. Avoid: Restrict baseline metrics strictly to the pre-assignment time window.
- Mistake: Interpreting aggregated experiment lifts without inspecting traffic splits across user segments. Avoid: Verify sample ratio mismatch and evaluate stratified cohort metrics.
- Mistake: Relying on asymptotic Student t-tests for heavy-tailed or extreme-revenue metrics. Avoid: Run non-parametric permutation tests or bootstrap resamples for skewed data.
- Mistake: Applying Difference-in-Differences without validating historical pre-intervention trajectories. Avoid: Test pre-period parallel trends across multiple pre-treatment timestamps.
FAQs
- How does CUPED reduce required experiment sample size? Because sample size is proportional to metric variance, reducing variance by $(1 - \rho^2)$ directly reduces the sample size and test duration needed for target power.
- What causes Simpson's Paradox in online A/B testing? It occurs when unequal traffic allocation ratios interact with heterogeneous conversion rates across confounding segments like mobile versus desktop users.
- When is a permutation test preferred over the Mann-Whitney U test? Permutation tests directly evaluate differences in means on the original scale, whereas Mann-Whitney evaluates rank differences that do not test expected value differences.
- What happens if parallel trends fail in a Difference-in-Differences design? The interaction coefficient will conflate existing baseline trend divergence with the true treatment effect, resulting in a biased and invalid causal estimate.