This lesson on Sample Size & Minimum Detectable Effect - The Napkin Version is hands-on and example-driven. Calculate required sample sizes per variant for A/B tests using the napkin power analysis formula (16 * variance / delta^2). You will connect significance level (Alpha), statistical power (1 - Beta), metric variance, and Minimum Detectable Effect (Delta) directly to product trade-offs in analytics interviews.
What You'll Be Able To Do
- Calculate sample size per variant using the napkin formula 16 * sigma^2 / delta^2
- Translate business risk tolerance into appropriate Alpha and Beta error thresholds
- Estimate metric variance from historical telemetry or pre-experiment A/A testing
- Set Minimum Detectable Effect thresholds based on practical business significance
- Evaluate trade-offs between test run time, detectable effect size, and statistical power
Detailed Concept Walkthrough
1. The Core Power Analysis Formula
Sample size estimation determines the minimum observation count needed per variant to detect a real treatment effect while controlling error rates. The napkin formula approximates exact power analysis using industry-standard statistical defaults.
- Mechanism: The classic sample size equation balances critical z-scores against baseline variance and expected lift: n = 2 * (Z_alpha/2 + Z_beta)^2 * sigma^2 / delta^2.
- Under the Hood: For standard industry defaults of alpha = 0.05 (Z = 1.96) and power = 0.80 (Z = 0.84), the critical term evaluates to 2 * (1.96 + 0.84)^2 = 15.68, which rounds to 16.
- Best Practice: In interview product cases, leverage the napkin approximation n = 16 * sigma^2 / delta^2 for fast mental math before writing code.
# Napkin formula implementation for standard 5% alpha and 80% power
def napkin_sample_size(variance: float, delta: float) -> int:
import math
# 16 * sigma^2 / delta^2 per test variant
return math.ceil(16 * variance / (delta ** 2))
# Example: baseline conversion p=0.10, desired absolute MDE=0.01
var_cr = 0.10 * (1 - 0.10) # 0.09
n_needed = napkin_sample_size(variance=var_cr, delta=0.01) # 14,400 per variant
Key Takeaway: Sample size scales linearly with metric variance and inversely with the square of the detectable effect size.
2. Alpha and Type I False Positive Error
Alpha represents the significance level, which is the risk of detecting a statistically significant difference when no true effect exists (a false positive).
- Mechanism: Lowering alpha from 0.05 to 0.01 shifts the critical normal threshold Z_alpha/2 outward from 1.96 to 2.58, widening the confidence interval.
- Under the Hood: A stricter false positive threshold requires stronger empirical evidence to reject the null hypothesis, directly driving up the required sample size multiplier.
- Best Practice: Maintain alpha at 0.05 for iterative product experiments, but lower it to 0.01 for high-risk irreversible changes such as removing core UI features or altering billing logic.
import scipy.stats as stats
# Calculate two-tailed critical value Z for given alpha
def get_z_alpha(alpha: float) -> float:
return stats.norm.ppf(1 - alpha / 2)
z_05 = get_z_alpha(0.05) # 1.95996 (~1.96)
z_01 = get_z_alpha(0.01) # 2.57583 (~2.58)
Key Takeaway: Tighter false positive tolerances require larger critical Z-scores and substantially increase required sample size.
3. Beta, Power, and Type II False Negative Error
Beta is the probability of failing to detect a real effect (false negative), while Power (1 - Beta) measures the experiment's ability to successfully identify an actual lift.
- Mechanism: Standard product experimentation sets power to 80% (beta = 0.20, Z_beta = 0.84), accepting a 20% risk of missing a true difference.
- Under the Hood: Increasing power to 90% (beta = 0.10, Z_beta = 1.28) increases the combined multiplier (Z_alpha/2 + Z_beta)^2 from ~7.84 to ~10.50, requiring roughly 33% more data.
- Best Practice: Choose 80% power for standard UI iterations; increase to 90% only when failing to detect a winning feature imposes high strategic or engineering opportunity costs.
# Compute exact sample size multiplier for custom alpha and power
def calculate_multiplier(alpha: float, power: float) -> float:
z_a = stats.norm.ppf(1 - alpha / 2)
z_b = stats.norm.ppf(power)
return 2 * ((z_a + z_b) ** 2)
mult_80 = calculate_multiplier(0.05, 0.80) # ~15.68
mult_90 = calculate_multiplier(0.05, 0.90) # ~21.02
Key Takeaway: Increasing statistical power protects against false negatives but requires substantially larger sample sizes per variant.
4. Variance Estimation and Minimum Detectable Effect
Variance measures metric dispersion, while Delta (MDE) represents the smallest treatment lift that justifies product rollout from a business ROI perspective.
- Mechanism: For binary metrics (CTR, conversion rate), variance is calculated analytically as p * (1 - p); for continuous metrics (revenue per user), variance is calculated from empirical baseline distributions.
- Under the Hood: If no historical telemetry exists, teams run an A/A test to measure baseline metric variance and ensure false positive rates behave as expected under the null distribution.
- Best Practice: Ground delta in practical business significance (e.g., revenue gains covering rollout cost) rather than statistical convenience, because halving delta quadruples needed sample size.
# Calculate exact sample size for proportion metric with relative MDE
def proportion_sample_size(p_baseline: float, relative_mde: float, alpha=0.05, power=0.80) -> int:
import math
delta = p_baseline * relative_mde
variance = p_baseline * (1 - p_baseline)
mult = calculate_multiplier(alpha, power)
return math.ceil(mult * variance / (delta ** 2))
# Baseline 5% CTR, target 10% relative lift (delta = 0.005)
n_exact = proportion_sample_size(0.05, 0.10) # 30,123 users per variant
Key Takeaway: Halving the Minimum Detectable Effect quadruples required sample size due to the inverse squared relationship.
Topics Covered in Sample Size & Minimum Detectable Effect - The Napkin Version
- Interview Framing (0:40 - 2:03) — Explains why sample size estimation matters in data science product case interviews.
- Power Analysis Formula (2:03 - 3:20) — Introduces the power analysis equation and its four primary mathematical components.
- Alpha and False Positives (3:20 - 4:15) — Covers significance levels, Type I error rates, and their effect on sample size.
- Beta and Statistical Power (4:15 - 4:56) — Details Type II errors, power (1 - Beta), and the cost of detecting true lifts.
- Metric Variance (4:56 - 5:36) — Demonstrates estimating metric dispersion from historical records and A/A tests.
- Delta and MDE (5:36 - 7:38) — Connects Minimum Detectable Effect to business decision-making and practical significance.
Stats & Product Analytics for Analysts Cheat Sheet
-
Napkin Sample Size— Approximates required sample size per variantn = 16 * variance / (delta ** 2) -
Proportion Variance— Calculates variance for binary metricsvariance = p * (1 - p) -
Absolute MDE Calculation— Converts relative lift to absolute deltadelta = baseline_metric * relative_lift -
Alpha Critical Z-Score— Computes two-tailed standard normal critical thresholdz_alpha = stats.norm.ppf(1 - alpha / 2) -
Power Critical Z-Score— Computes one-tailed standard normal power thresholdz_power = stats.norm.ppf(power) -
Exact Multiplier Formula— Calculates sample size coefficient from Z-scoresmultiplier = 2 * ((z_alpha + z_power) ** 2)
Comparison Table
| Parameter | Statistical Role | Sample Size Impact |
|---|---|---|
| Alpha (Type I) | False positive error rate | Lower alpha increases sample size |
| Power (1 - Beta) | Probability of detecting effect | Higher power increases sample size |
| Variance (sigma^2) | Noise in baseline metric | Higher variance increases sample size |
| MDE (delta) | Minimum meaningful business effect | Smaller delta quadratically increases size |
Common Pitfalls
- Mistake: Setting MDE arbitrarily small without business rationale. Avoid: Align MDE with economic viability because halving delta quadruples required sample size.
- Mistake: Confusing relative percentage lift with absolute delta. Avoid: Multiply baseline metric by relative lift before inserting delta into sample size equations.
- Mistake: Assuming zero baseline variance for continuous metrics. Avoid: Compute empirical variance using historical user data or a pre-experiment A/A test.
- Mistake: Using the napkin multiplier 16 when alpha or power change. Avoid: Recalculate 2 * (Z_alpha/2 + Z_beta)^2 whenever departing from alpha 0.05 and power 0.80.
FAQs
- Where does the constant 16 come from in the napkin formula? For alpha = 0.05 (Z = 1.96) and power = 0.80 (Z = 0.84), 2 * (1.96 + 0.84)^2 equals 15.68, which rounds to 16.
- What is the difference between statistical significance and practical significance? Statistical significance confirms a difference is unlikely due to random noise, while practical significance confirms the effect size is large enough to matter for business ROI.
- How can I estimate variance if no historical baseline data is available? Run a pre-experiment A/A test on the target user population to measure empirical variance and confirm false positive stability.