This lesson on Regularization — L1, L2, Early Stopping is hands-on and example-driven. You will be able to apply Ridge Regression (L2 regularization) to linear models to effectively manage the bias-variance trade-off. You will learn how the L2 penalty shrinks coefficients and how to tune the regularization hyperparameter (lambda) using cross-validation.
What You'll Be Able To Do
- Define the bias-variance trade-off in the context of model generalization.
- Formulate the penalized cost function for Ridge Regression.
- Calculate the effect of the L2 penalty on model coefficients (shrinkage).
- Select the optimal regularization strength (lambda) using cross-validation.
- Apply L2 regularization principles to both continuous and discrete predictors.
Detailed Concept Walkthrough
1. Bias-Variance Trade-off
Overfitting occurs when a model fits the training data too closely, resulting in high variance and poor generalization to new data. Ridge Regression intentionally introduces a small amount of bias to significantly reduce this variance, improving long-term prediction stability.
- Mechanism: Least Squares minimizes only the training error, often leading to overly complex models with large coefficients that are highly sensitive to small changes in the input data.
- Best Practice: When training data is small or noisy, high variance is common. Regularization is the preferred technique over simply gathering more data, especially when data acquisition is costly.
- Under the Hood: The goal is not to minimize training error (Cost) but to minimize the expected prediction error (Cost + Penalty), which forces the model to prioritize simplicity over perfect training fit.
import numpy as np
from sklearn.linear_model import LinearRegression, Ridge
# Simulate high variance data (small dataset)
X = np.array([[1], [2], [10], [11]])
y = np.array([1, 2, 10, 11])
# 1. High Variance (Least Squares)
ls_model = LinearRegression().fit(X, y)
# 2. Regularized (Ridge)
ridge_model = Ridge(alpha=1.0).fit(X, y)
print(f"LS Coef: {ls_model.coef_[0]:.2f}")
print(f"Ridge Coef (alpha=1): {ridge_model.coef_[0]:.2f}") # Coefficient is shrunk
Key Takeaway: Better generalization requires sacrificing a perfect training fit (introducing bias) to gain stability (reducing variance).
2. L2 Regularization Cost Function
Ridge Regression minimizes the standard Sum of Squared Residuals (SSR) plus an L2 penalty term proportional to the square of the coefficients. This penalty discourages large coefficient magnitudes.
- Mechanism: The cost function is $J(\theta) = \text{SSR} + \lambda \sum_{j=1}^p \theta_j^2$. The $\sum \theta_j^2$ term is the L2 norm squared, which penalizes large weights quadratically.
- Execution Flow: During optimization, the algorithm must balance minimizing the fit error (SSR) and minimizing the size of the coefficients (L2 penalty).
- Under the Hood: Since the penalty is based on the square of the coefficient, it forces all coefficients to shrink towards zero, but rarely exactly to zero, maintaining all features in the model.
# Conceptual Cost Function (Python pseudo-code)
# Cost = Sum of Squared Residuals (SSR)
def calculate_ssr(y_true, y_pred):
return sum((y_true - y_pred)**2)
# Penalty = Lambda * L2 Norm Squared
def calculate_l2_penalty(lambda_val, coefficients):
l2_norm_sq = sum(c**2 for c in coefficients)
return lambda_val * l2_norm_sq
# Ridge Cost = SSR + Penalty
def ridge_cost(y_true, y_pred, lambda_val, coefficients):
return calculate_ssr(y_true, y_pred) + calculate_l2_penalty(lambda_val, coefficients)
Key Takeaway: The L2 penalty adds a cost for complexity, forcing the model to prefer smaller, less sensitive coefficients.
3. Coefficient Shrinkage and Desensitization
The L2 penalty shrinks the magnitude of the model coefficients towards zero, reducing the model's sensitivity to changes in the predictor variables. This process is also called desensitization.
- Mechanism: A steep slope (large coefficient) implies high sensitivity: a small change in X results in a large change in Y. The L2 penalty targets these large coefficients directly, reducing their magnitude.
- Under the Hood: The penalty term $\lambda \sum \theta_j^2$ grows rapidly as coefficients increase, making large coefficients prohibitively expensive in the overall cost function, thus forcing shrinkage.
- Best Practice: Shrinkage applies uniformly across all predictor types, including continuous variables and coefficients representing differences between discrete groups (e.g., diet type).
Key Takeaway: Shrinking coefficients reduces model sensitivity, leading to smoother prediction surfaces and better generalization.
4. Tuning the Regularization Parameter (λ)
The hyperparameter $\lambda$ controls the strength of the regularization penalty. Selecting the optimal $\lambda$ is crucial for balancing bias and variance.
- Mechanism: If $\lambda=0$, Ridge Regression reverts to standard Least Squares. As $\lambda$ increases, the penalty dominates, shrinking coefficients asymptotically towards zero.
- Execution Flow: Optimal $\lambda$ is selected using cross-validation (typically 10-fold). The model is trained across a range of $\lambda$ values, and the value minimizing the validation error is chosen.
- Best Practice: $\lambda$ is external to the model's training optimization; it must be tuned separately using validation data to prevent overfitting the regularization strength itself.
from sklearn.model_selection import GridSearchCV
from sklearn.linear_model import Ridge
# Define a range of lambda (alpha) values to test
param_grid = {'alpha': [0.1, 1.0, 3.0, 10.0, 100.0]}
# Use 10-fold cross-validation (cv=10) to find the best alpha
ridge_search = GridSearchCV(Ridge(), param_grid, cv=10)
# Fit the search object to the data
# ridge_search.fit(X_train, y_train)
# print(f"Optimal Lambda: {ridge_search.best_params_['alpha']}")
Key Takeaway: Lambda is the dial for complexity; tune it via cross-validation to find the sweet spot between bias and variance.
Topics Covered in Regularization — L1, L2, Early Stopping
- Regularization Context (0:00 - 1:15) — Regularization is introduced as a method to improve model generalization by managing the bias-variance trade-off.
- Addressing High Variance (1:17 - 3:40) — Least Squares often overfits small datasets, requiring regularization to stabilize predictions and reduce variance.
- L2 Cost Function (3:43 - 6:05) — The Ridge cost function minimizes the sum of squared residuals plus the L2 penalty term, $\lambda * (\text{Slope})^2$.
- Coefficient Shrinkage (6:07 - 7:20) — The L2 penalty reduces the magnitude of coefficients, decreasing the model's sensitivity to inputs (desensitization).
- Tuning Lambda (λ) (7:22 - 8:34) — Optimal $\lambda$ is found using cross-validation to select the regularization strength that minimizes validation error.
- Discrete Variables (8:36 - 10:50) — Ridge principles apply to discrete variables by shrinking the coefficient that represents the difference between groups.
- Logistic Regression (11:02 - 11:20) — L2 regularization can be applied to non-linear models like Logistic Regression to stabilize predictions.
ML Foundations Cheat Sheet
-
Ridge Regression— L2 regularization to manage variance and shrink coefficientsfrom sklearn.linear_model import Ridge model = Ridge(alpha=1.0) -
Lambda (λ)— Hyperparameter controlling the severity of the L2 penaltymodel = Ridge(alpha=3) -
L2 Penalty— Sum of squared coefficients added to the cost functioncost + lambda * sum(coef**2) -
Cross-Validation— Method to select optimal λ by minimizing validation errorGridSearchCV(Ridge(), params) -
Coefficient Shrinkage— Reduces coefficient magnitude towards zero (desensitization)
Comparison Table
| Least Squares | Ridge Regression (L2) | Lasso Regression (L1) |
|---|---|---|
| Minimizes SSR only | Minimizes SSR + L2 penalty | Minimizes SSR + L1 penalty |
| High variance/overfitting | Shrinks coefficients towards zero | Shrinks coefficients to zero |
| No bias introduced | Introduces small bias | Performs feature selection |
| λ is zero | λ is tuned via CV | λ is tuned via CV |
Common Pitfalls
- Mistake: Assuming λ=0 is always best because it minimizes training error. Avoid: Use cross-validation to find λ that minimizes test error.
- Mistake: Applying Ridge without standardizing features first. Avoid: Always scale features so the penalty applies equally to all coefficients.
- Mistake: Tuning λ on the training set itself. Avoid: Use a separate validation set or cross-validation for hyperparameter tuning.
FAQs
- Why is the penalty called L2? It uses the L2 norm (Euclidean distance) of the coefficient vector in the penalty term, which involves squaring the coefficient magnitudes.
- Does Ridge Regression set coefficients exactly to zero? No, the L2 penalty shrinks them asymptotically towards zero but keeps all features in the model, unlike L1 (Lasso).
- Can I use L2 regularization with non-linear models? Yes, L2 regularization can be applied to non-linear models like Logistic Regression to stabilize predictions and shrink estimated coefficients.
- What happens if lambda (λ) is too large? The model becomes overly biased (underfit), as coefficients are shrunk too close to zero, losing almost all predictive power.