Loading content...
Loading content...
Transform raw tables into actionable business intelligence. Master central tendency (mean, median, mode), quantify dispersion (range, variance, standard deviation), analyze quartile distributions (Q1, Q2, Q3, IQR), and leverage describe() to rapidly extract the statistical story behind your data.
Numerical summaries that describe essential data characteristics
Consider this simple salary column: 30000, 32000, 35000, 40000, 45000. What can an analyst learn without manually inspecting every individual row?
What is a typical salary?
How dispersed are values?
What are min & max?
Where do most peers sit?
Are there unusual spikes?
Summing values divided by observation count
The mean is the most common measure of center. In Pandas:
df["Salary"].mean()Sum of values ÷ number of values. For 30k, 32k, 35k, 40k, 45k, the mean is ₹36,400.
30k, 32k, 35k, 40k, 350kReplace 45k with 350k: the mean jumps to ₹97,400! Not a single typical employee earns anywhere near ₹97,400.
Position-based center that resists extreme outliers
Sort your data in ascending order; the median is the exact physical midpoint:
Exact middle element is the 3rd value: ₹35,000.
Average of the two middle elements (32k + 36k) / 2 = ₹34,000.
Vectorized, highly efficient, and automatically ignores NaNs.
Indispensable for categorical columns and product popularity
The mode is the observation that appears with the highest frequency:
["Mumbai", "Delhi", "Mumbai", "Pune", "Mumbai"]"Mumbai" appears 3 times. Mode = "Mumbai".
df["City"].mode()Returns a Series (because datasets can be bimodal with ties for 1st place).
Establishing boundaries and overall data boundaries
Smallest observed value in the column.
Largest observed value in the column.
Total distance spanned between lowest and highest point.
Dividing ordered data into four quarters
• Q1 (25th percentile): df["Salary"].quantile(0.25)
• Q2 (50th percentile / Median): df["Salary"].quantile(0.50)
• Q3 (75th percentile): df["Salary"].quantile(0.75)
IQR = Q3 - Q1Measures the spread of the middle 50% of your data. Completely immune to extreme minimums and maximums, serving as the statistical backbone for outlier detection.
Quantifying how observations scatter around the mean
Why isn’t the mean enough? Compare these two groups of test scores:
48, 49, 50, 51, 52• Mean: 50.0
• Standard Deviation: 1.58
• Highly consistent; scores are clustered tightly near 50.
20, 35, 50, 65, 80• Mean: 50.0 (Identical mean!)
• Standard Deviation: 23.18
• Highly volatile; scores are widely scattered from the center.
df["Salary"].var() → Sample variance (in squared units).df["Salary"].std() → Standard deviation (in original units, e.g. Rupees or Years). A larger standard deviation indicates greater variability.Extracting an instant 8-point numerical overview of every column
Calling df.describe() computes all essential metrics in a single line:
| Metric | Meaning | Analyst Takeaway |
|---|---|---|
count | Total non-null rows | Quick check for missing data |
mean | Arithmetic average | General center |
std | Sample standard deviation | Spread around average |
min | Lowest observed value | Absolute floor |
25% | First quartile (Q1) | 25% of data sits below here |
50% | Median (Q2) | Exact physical middle |
75% | Third quartile (Q3) | 75% of data sits below here |
max | Highest observed value | Absolute ceiling |
Generate statistics and evaluate whether Salary has an extreme right-skew
Observe how an outlier shifts the mean by 70% while leaving median untouched
How describe() summarizes text columns
When invoked on non-numeric columns like df["City"].describe(), Pandas produces:
Total non-null string entries.
Count of distinct categories/values.
The mode (most frequent string).
Number of times the "top" string occurred.
Synthesizing statistics across multiple business dimensions
| Employee | Department | Age | Salary | Performance |
|---|---|---|---|---|
| Amit | Sales | 24 | ₹42,000 | 78 |
| Priya | HR | 27 | ₹50,000 | 85 |
| Rahul | Sales | 25 | ₹45,000 | 82 |
| Neha | IT | 29 | ₹60,000 | 91 |
| Karan | Sales | 31 | ₹48,000 | 80 |
| Riya | IT | 26 | ₹65,000 | 88 |
| Vikas | HR | 35 | ₹120,000 | 86 |
Numerical summaries are step one in an ongoing discovery cycle
The mathematical numbers that quantify center, spread, percentiles, and boundaries.
Answers:"What are the numbers?"
The investigative holistic process combining descriptive statistics, charts, scatter plots, correlations, and business questions.
Answers:"What business patterns do these numbers reveal?"
Analyze product pricing, sales volume dispersion, and extract business insights
You are auditing an e-commerce catalog with 8 products. Generate the full descriptive statistics table, compare mean vs median price, calculate range and IQR, and examine UnitsSold standard deviation:
Pitfalls made when interpreting descriptive statistics
In right-skewed data, the mean is pulled upwards. Always compare mean with median to check for skewness.
Range only measures the difference between the two most extreme points (Max - Min). Standard deviation accounts for how all points scatter around the mean.
Descriptive statistics are merely the numerical summary phase. Full EDA requires visual exploration, cross-tabulation, and investigating business causality.
Q1 and Q3 mark the 25th and 75th percentiles (the middle 50% boundary). Min and Max represent the absolute single lowest and highest points.
Validate your descriptive statistics mastery
Validate your understanding of mean, median, mode, range, quartiles, IQR, standard deviation, and describe().
Why is the sample median often preferred over the sample mean when analyzing income or real-estate prices?
Skills you have mastered in Descriptive Statistics
.mean(), .median(), and .mode()..min(), .max(), and Range..quantile() and determine the IQR..var()) and standard deviation (.std()).df.describe().