Back to Modeling: Regression and Beyond

Linear Regression — The Math, Not Just the API

What `LinearRegression().fit()` actually does, and the assumptions that make it valid. FIND_VIDEO: search 'linear regression OLS assumptions' — recommended channel: StatQuest / 3Blue1Brown. Aim for 11 min or under.

27 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Core Regression Steps — Regression analysis involves fitting a line, quantifying performance, and testing statistical significance.
  2. Least Squares Method — The best-fit line minimizes the Sum of Squared Residuals (SSR), which are the errors between data and the line.
  3. Model Parameters — The fitted line is defined by the estimated slope and Y-intercept, which determine its predictive power.
  4. Baseline Variation — SS Mean establishes the total variation in Y that must be explained, calculated around the average Y value.
  5. Residual Variation — SS Fit quantifies the unexplained variation remaining after the least squares line has been fitted to the data.
  6. R-Squared Formula — R-squared is calculated as the ratio of explained variation to total variation, showing goodness-of-fit.
  7. Boundary Cases — R-squared ranges from 0% (no fit) to 100% (perfect fit) and applies universally across model types.
  8. Multiple Predictors — Linear regression extends to multiple variables by fitting a plane or hyperplane in higher dimensions.
PDF notes

Frequently asked questions

Why do we square the residuals instead of just summing the absolute values?

Squaring penalizes larger errors more heavily, ensuring the line is closer to all points. It also makes the error function differentiable, which is necessary for optimization algorithms.

Does R-squared work for non-linear models?

Yes, the R-squared calculation based on SS Mean and SS Fit is universal for quantifying the goodness-of-fit of any model equation, provided it minimizes squared error.

What does a non-zero slope imply?

A non-zero slope suggests that the independent variable (X) has predictive power over the dependent variable (Y), meaning the fitted line is better than just using the mean.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.