Back to Supervised Learning

Regularization — L1, L2, Early Stopping

The three techniques that prevent overfitting. When and how to use each. FIND_VIDEO: search 'regularization L1 L2 dropout machine learning' — recommended channel: StatQuest. Aim for 10 min or under.

20 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Regularization Context — Regularization is introduced as a method to improve model generalization by managing the bias-variance trade-off.
  2. Addressing High Variance — Least Squares often overfits small datasets, requiring regularization to stabilize predictions and reduce variance.
  3. L2 Cost Function — The Ridge cost function minimizes the sum of squared residuals plus the L2 penalty term, $\lambda * (\text{Slope})^2$.
  4. Coefficient Shrinkage — The L2 penalty reduces the magnitude of coefficients, decreasing the model's sensitivity to inputs (desensitization).
  5. Tuning Lambda (λ) — Optimal $\lambda$ is found using cross-validation to select the regularization strength that minimizes validation error.
  6. Discrete Variables — Ridge principles apply to discrete variables by shrinking the coefficient that represents the difference between groups.
  7. Logistic Regression — L2 regularization can be applied to non-linear models like Logistic Regression to stabilize predictions.
PDF notes

Frequently asked questions

Why is the penalty called L2?

It uses the L2 norm (Euclidean distance) of the coefficient vector in the penalty term, which involves squaring the coefficient magnitudes.

Does Ridge Regression set coefficients exactly to zero?

No, the L2 penalty shrinks them asymptotically towards zero but keeps all features in the model, unlike L1 (Lasso).

Can I use L2 regularization with non-linear models?

Yes, L2 regularization can be applied to non-linear models like Logistic Regression to stabilize predictions and shrink estimated coefficients.

What happens if lambda (λ) is too large?

The model becomes overly biased (underfit), as coefficients are shrunk too close to zero, losing almost all predictive power.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.