Back to PANDAS FOR DATA ANALYTICS

Groupby & Aggregations

Learn to summarize and aggregate data efficiently using groupby—one of the most powerful operations in pandas.

11 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Groupby Concept — The `groupby` function segments data based on column values to allow aggregate functions to run on those segments.
  2. GroupBy Object — Calling `df.groupby()` returns a DataFrameGroupBy object, which is an intermediate step before calculation.
  3. Applying .mean() — Applying `.mean()` calculates the average for all numeric columns within each group, automatically excluding string columns.
  4. Count, Min, Max — Standard aggregations like `.count()`, `.min()`, and `.max()` are demonstrated, noting that min/max work on strings alphabetically.
  5. Using .agg() — The `.agg()` function allows custom aggregation by passing a dictionary mapping columns to a list of desired functions.
  6. Multiple Grouping — Data can be grouped by multiple columns by passing a list of column names to `groupby()`.
  7. .describe() Shortcut — The `.describe()` function provides a quick, generalized overview of multiple aggregate statistics for all numeric columns.
PDF notes

Frequently asked questions

Why did `.groupby()` return an object instead of a DataFrame?

`groupby()` creates an intermediate `DataFrameGroupBy` object; you must call an aggregation function (like `.mean()` or `.agg()`) on it to execute the calculation and return a DataFrame.

How do `.min()` and `.max()` handle string columns?

They compare strings alphabetically. `.min()` returns the value that comes earliest in the alphabet, and `.max()` returns the value that comes latest.

Can I group by more than one column?

Yes, pass a list of column names to the `groupby()` function, e.g., `df.groupby(['base flavor', 'liked'])`.

What is the benefit of `.describe()` over simple aggregations?

`.describe()` is a shortcut that returns seven common statistics (count, mean, std, min, quartiles, max) for all numeric columns simultaneously, providing a quick overview.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.