Apply a Custom Function
Categorize employees into salary bands using a custom function. Available Data: DataFrame df with columns: name,...
Select, group, merge and derive columns on a table in Python, without loops.
A pandas DataFrame is a labelled, two-dimensional table in Python, and most analysis with it comes down to four moves: select the rows and columns you need, group and aggregate them, merge in another table, and derive new columns without looping. Doing that vectorised, with whole-column operations instead of a Python for loop, is what keeps the code both short and fast. These questions run real pandas in the browser against seeded DataFrames, so the output you check is the actual frame.
Categorize employees into salary bands using a custom function. Available Data: DataFrame df with columns: name,...
Given a 2D numpy array matrix (rows = observations, columns = features), compute the z-score-standardized version using...
Given a 2D numpy array points of shape (N, 2) representing N 2D points, compute the N×N pairwise Euclidean distance...
Given a 1D numpy array of float readings, extract a sorted array of all readings that are: Greater than 50, AND Less...
Given a 1D numpy array values of length N, compute three series and stack them into a 2D array of shape (N, 3): Column...
A mobile banking app's onboarding funnel needs a step-conversion breakdown for the weekly product review. You're given...
You're given a small event log of user activity: each row is one (user, month) pair in which that user was active,...
Lowercase the plan codes in df and store the updated frame in result.
Group df by load_date, summing mrr into partition_mrr with row counts n_rows. Store in result.
FAQ
Short answers to what people get stuck on most. Tap a question to expand it.
Aggregations skip NaN by default, so a group still returns a total from its non-missing values.
df.groupby("city")["sales"].sum() # all-NaN group -> 0.0
df.groupby("city")["sales"].sum(min_count=1) # all-NaN group -> NaNloc selects by label, iloc selects by integer position.
df.loc[df["sales"] > 100, "city"] # label + boolean mask
df.iloc[0:2, 0] # first two rows, first columnUse merge, which works like a SQL join. Choose how: inner, left, right or outer.
orders.merge(
customers,
on="customer_id",
how="left",
validate="many_to_one",
)Track your streak & earn XP
Sign up free to unlock your progress dashboard, daily streaks, and leaderboard ranking.
10 role-based tests for Data Analysts, SQL Developers & Data Engineers with instant scorecards.