Remove Duplicate Rows
The dataset has duplicate rows. Remove them and keep only the first occurrence. Available Data: DataFrame df with...
The dataset has duplicate rows. Remove them and keep only the first occurrence. Available Data: DataFrame df with...
Count the number of missing (null/NaN) values in each column. Available Data: DataFrame df with columns: name, email,...
Remove all rows that contain any missing values. Available Data: DataFrame df with columns: name, email, age, salary...
Fill missing salary values with the mean salary of the existing values. Available Data: DataFrame df with columns:...
Given a DataFrame with numeric data stored as strings, convert the price column to float type. Available Data:...
Replace the department name 'HR' with 'Human Resources' in the employee DataFrame. Available Data: DataFrame df with...
Remove the city column from the employee DataFrame. Available Data: DataFrame df with columns: name, department,...
Clean the messy dataset by removing both null values and duplicate rows. Available Data: DataFrame df with columns:...
Drop rows only where the email column is null. Keep rows with null values in other columns. Available Data: DataFrame...
Clip salary values to be between 60,000 and 100,000. Values below 60k should become 60k, values above 100k should...
Calculate the average salary for each department. Available Data: DataFrame df with columns: name, department, salary,...
For each department, find the minimum, maximum, and mean salary. Available Data: DataFrame df with columns: name,...
Count the number of employees in each city. Available Data: DataFrame df with columns: name, department, salary, age,...
Merge the orders DataFrame with the customers DataFrame to see customer names with their orders. Available Data:...
Perform a left merge to see all customers, even those without orders. Available Data: orders: order_id, customer_id,...
Track your streak & earn XP
Sign up free to unlock your progress dashboard, daily streaks, and leaderboard ranking.
10 role-based tests for Data Analysts, SQL Developers & Data Engineers with instant scorecards.
Select, group, merge and derive columns on a table in Python, without loops.
A pandas DataFrame is a labelled, two-dimensional table in Python, and most analysis with it comes down to four moves: select the rows and columns you need, group and aggregate them, merge in another table, and derive new columns without looping. Doing that vectorised, with whole-column operations instead of a Python for loop, is what keeps the code both short and fast. These questions run real pandas in the browser against seeded DataFrames, so the output you check is the actual frame.
FAQ
Short answers to what people get stuck on most. Tap a question to expand it.
Aggregations skip NaN by default, so a group still returns a total from its non-missing values.
df.groupby("city")["sales"].sum() # all-NaN group -> 0.0
df.groupby("city")["sales"].sum(min_count=1) # all-NaN group -> NaNloc selects by label, iloc selects by integer position.
Use merge, which works like a SQL join. Choose how: inner, left, right or outer.
df.loc[df["sales"] > 100, "city"] # label + boolean mask
df.iloc[0:2, 0] # first two rows, first columnorders.merge(
customers,
on="customer_id",
how="left",
validate="many_to_one",
)