Python for Data Analytics • Module 06 • Scatter Plots

Scatter Plots: Correlation, Clusters & Outlier Analysis

Master bivariate correlation, cluster discovery, and outlier analysis using Scatter Plots in Python for data analytics. Learn explicit Figure/Axes ax.scatter(), alpha transparency against overplotting, bubble charts with multivariate size/color encodings, trendlines with np.polyfit(), and Anscombe's quartet.

Estimated Time: 45 Minutes
Level: Beginner to Intermediate
Track: Python 3.12+ / Matplotlib / Pandas / NumPy
Interface: Explicit Figure/Axes (fig, ax = plt.subplots())
01

The Analytical Foundation of Scatter Plots

Bivariate spatial encoding, Cleveland's perception hierarchy, and Anscombe's revelation

In data visualization science, 2D position along common aligned scales $(x, y)$ provides the highest quantitative perceptual decoding accuracy among all visual encodings (superior to length, area, angle, or color hue). Unlike line charts, which imply temporal progression between adjacent points, a scatter plot treats each point as an independent observation, rendering bivariate co-variation without manufactured continuity.

The Exploratory Discovery Spectrum of Scatter Plots
Linear Co-Variation
Positive or negative slope indicating direct or inverse mathematical relationship.
→
Natural Clustering
Discrete customer segments, price tiers, or operational efficiency groups.
→
Heteroscedasticity
Fan-shaped spread where variance expands or contracts as X increases.
→
Leverage Outliers
Extreme anomalies that pull regression slopes and distort correlation metrics.
Why Summary Statistics Lie (Anscombe's Warning)
Never rely solely on summary statistics like mean, standard deviation, and Pearson $r$. Four completely distinct datasets can yield identical summary numbers (e.g., $r = 0.816$), yet one is a clean line, one is a parabolic curve, one is a line with an extreme vertical outlier, and one is a vertical cluster with a high-leverage outlier. Always visualize bivariate data with a scatter plot before making business claims.
02

Explicit Figure/Axes & ax.scatter() Anatomy

Signature parameters, marker scaling in points squared ($pt^2$), and PathCollection performance

In modern Matplotlib 2026, we invoke ax.scatter() on an explicit Axes instance. Under the hood, ax.scatter() returns a PathCollection object, which allows individual points to have distinct sizes, edge colors, and colormap mapping values.

ParameterUnits / TypesAnalytical Purpose & Critical Gotchas
x, y1D array-like, SeriesCoordinates of the observations. Must have identical lengths.
s (Size)Scalar or Array of floatsPoints squared ($pt^2$): Area of the marker symbol. Do NOT pass raw radius; pass area ($s \propto r^2$) to avoid geometric perceptual distortion.
c (Color)Color string or Array of valuesSingle color (e.g. '#10b981') or array of continuous values mapped to a colormap.
alphaFloat (0.0 to 1.0)Translucency. Vital for combating overplotting. Use 0.2–0.5 for dense datasets.
cmapMatplotlib ColormapColormap name (e.g. 'viridis', 'plasma') applied when c is numeric.
edgecolorsColor or 'none'Border stroke around points. White edges ('#ffffff') help distinguish overlapping bubbles.
Python • Explicit Figure/Axes Scatter Plot
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd

# 1. Initialize explicit Figure and Axes (2026 standard)
fig, ax = plt.subplots(figsize=(10, 6), dpi=100)

# 2. Plot observations with area-scaled bubbles and alpha
scatter = ax.scatter(
    df['ad_spend'],
    df['revenue'],
    s=df['deal_size'] * 1.5,     # Area scaling in points^2
    c=df['profit_margin'],       # Continuous color mapping
    cmap='viridis',
    alpha=0.65,
    edgecolors='#ffffff',
    linewidths=0.75
)

# 3. Add continuous colorbar linked to scatter collection
cbar = fig.colorbar(scatter, ax=ax)
cbar.set_label('Profit Margin (%)', fontsize=10, weight='bold')

# 4. Polish typography and canvas
ax.set_title('Campaign Ad Spend vs Gross Merchandising Revenue', fontsize=13, weight='bold', pad=12)
ax.set_xlabel('Digital Ad Spend ($K USD)', fontsize=11, weight='bold')
ax.set_ylabel('Revenue ($K USD)', fontsize=11, weight='bold')
ax.grid(True, linestyle='--', alpha=0.3)

plt.tight_layout()
plt.show()
Interactive Lab 01

Real-Time Scatter Plot Simulator & Regression Tuner

Tweak marker area, alpha transparency, colormaps, regression trendlines, and quadrant crosshairs.

1-Click Presets:
28622516310240104173104135Digital Ad Spend ($K)Gross Merchandising Revenue ($K)
r = 0.963
Pearson Correlation
R² = 0.927
Variance Explained
m = 1.88
OLS Regression Slope
N = 31
Sample Points
Strong Positive
Analytical Diagnosis
Reactive Matplotlib Code Generator
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd

# 1. Explicit Figure and Axes (2026 standard)
fig, ax = plt.subplots(figsize=(10, 6), dpi=100)

# 2. Plot bivariate scatter points with alpha transparency
ax.scatter(
    df['Digital Ad Spend ($K)'],
    df['Gross Merchandising Revenue ($K)'],
    s=45,
    color='#10b981',
    alpha=0.65,
    edgecolors='#ffffff',
    linewidths=0.6,
    label='Observations (N=31)'
)

# 3. Fit and plot OLS Linear Trendline (NumPy)
slope, intercept = np.polyfit(df['Digital Ad Spend ($K)'], df['Gross Merchandising Revenue ($K)'], deg=1)
x_vals = np.linspace(df['Digital Ad Spend ($K)'].min(), df['Digital Ad Spend ($K)'].max(), 100)
ax.plot(x_vals, slope * x_vals + intercept, color='#f43f5e', lw=2.2, ls='--', label=f'OLS Trendline (r=0.963)')

# 5. Polish typography & gridlines
ax.set_title('Marketing Ad Spend vs Revenue ($K)', fontsize=14, weight='bold', pad=14)
ax.set_xlabel('Digital Ad Spend ($K)', fontsize=11, weight='bold')
ax.set_ylabel('Gross Merchandising Revenue ($K)', fontsize=11, weight='bold')
ax.grid(True, linestyle='--', alpha=0.3)
ax.legend(frameon=True, facecolor='#090f24', edgecolor='none')
plt.tight_layout()
plt.show()
03

The Overplotting Crisis & Alpha Translucency

Combating coordinate density collapse using alpha transparency, point sizing, and jittering

When working with real-world enterprise databases containing tens or hundreds of thousands of rows, points plotted with full opacity (alpha=1.0) overlap into a solid, impenetrable blob. The viewer cannot distinguish whether a region contains 5 observations or 5,000 observations.

Interactive Lab 02

Overplotting Simulator: Opaque Blob vs Alpha Density

Toggle between fully opaque overlap and translucent alpha to expose the hidden density core.

Key Rule: Whenever plotting $N > 500$ points, set alpha=0.15 to 0.35. For extreme datasets ($N > 500,000$), migrate to ax.hexbin() or 2D density contours.
04

Multivariate Bubble Charts (Size & Color Channels)

Encoding up to 5 continuous and categorical variables simultaneously

A standard scatter plot represents 2 variables $(X, Y)$. By introducing marker area ($s$) and colormap gradient ($c$), an analyst can communicate 4 continuous variables in a single coherent view without resorting to misleading 3D projections.

Interactive Lab 03

Multivariate Channel Allocation Lab

Select which analytical metric to map to Bubble Size and Continuous Color. (Practice starts blank).

05

Trendlines, Regression & Pearson Correlation (r)

Quantifying linear association and fitting Ordinary Least Squares with NumPy np.polyfit()

The Pearson correlation coefficient ($r$) measures the strength and direction of linear association, ranging from $-1.0$ (perfect negative slope) to $+1.0$ (perfect positive slope). In Python, calculate it directly with Pandas: df['x'].corr(df['y']).

Pearson Coefficient ($r$)Relationship StrengthBusiness Analytics Interpretation
$+0.80$ to $+1.00$Very Strong PositiveHigh predictability. As investment grows, outcome consistently rises.
$+0.40$ to $+0.79$Moderate PositiveDefinite directional correlation, but substantial variation remains unexplained.
$-0.39$ to $+0.39$Weak / NegligibleNo clear linear relationship. (Check for non-linear curves before concluding zero effect!)
$-0.80$ to $-1.00$Very Strong NegativeInverse relationship (e.g. Higher software price $\to$ Lower conversion rate).
06

Outlier Detection & Quadrant Segmentation

Dividing bivariate distributions into actionable 2x2 executive decision matrix quadrants

One of the highest-leverage applications of scatter plots in corporate analytics is 2x2 Quadrant Segmentation. By drawing crosshair benchmark lines at the median or target thresholds using ax.axvline() and ax.axhline(), the chart transforms from raw data into four prioritized business buckets.

Interactive Lab 05

Outlier Anomaly Annotation Builder

Type an anomaly reason label to annotate the high-leverage outlier with ax.annotate().

07

Seaborn Statistical Relational Plotting

Streamlining multivariate analysis with sns.scatterplot() and sns.regplot()

While Matplotlib gives raw foundational control, Seaborn dramatically simplifies multi-category scatter plots by automatically handling semantic groupings (hue), marker styles, and legend generation directly from DataFrame column names.

Python • Seaborn Relational Scatter & Regression
import seaborn as sns
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(10, 6), dpi=100)

# Automated hue, size, and style mappings
sns.scatterplot(
    data=df,
    x='cac',
    y='ltv',
    hue='customer_tier',     # Categorical color grouping
    size='user_seats',       # Bubble sizing
    sizes=(30, 200),
    palette='viridis',
    alpha=0.75,
    ax=ax
)

# Overlay linear regression trendline with 95% confidence interval
sns.regplot(
    data=df,
    x='cac',
    y='ltv',
    scatter=False,           # Don't duplicate points
    color='#f43f5e',
    ax=ax
)

ax.set_title('SaaS CAC vs LTV Cohort Health with 95% CI', fontsize=13, weight='bold')
plt.show()
08

Pandas Native Integration & Aggregations

Direct DataFrame plotting with df.plot.scatter() and group-level bivariate analysis

In production EDA, analysts frequently aggregate raw transactions using Pandas groupby() before passing them to a scatter plot. For instance, plotting Average Order Value vs Return Rate across 50 product categories.

Workflow StagePandas Code PatternAnalytical Purpose
Correlation Matrixdf[['spend', 'rev', 'margin']].corr()Identify candidate variable pairs for scatter plot investigation.
Group Aggregationagg_df = df.groupby('category').agg({'sales': 'mean', 'returns': 'mean'})Condense millions of transactions into category-level coordinates.
Direct Plottingagg_df.plot.scatter(x='sales', y='returns', s=50, ax=ax)Quick one-liner EDA directly from a Pandas DataFrame.
09

Production Incident Case Studies

Real-world analytical catastrophes: The Phantom Outlier Billionaire and Simpson's Paradox

Production Incident #1

The "Phantom Billionaire" Axis Compression Fiasco

A personal finance app created a customer net worth vs credit score scatter plot for 200,000 users. A single multi-billionaire client had a net worth of $4,500,000,000, while 99.9% of users had net worths under $500,000. On a linear scale, 199,999 users were crushed into a single flat line against the zero axis, making the entire visualization completely useless.

10

Decision Matrix & Anti-Patterns Checklist

Structured selection framework and critical executive visualization traps

Data Shape & Analytical GoalOptimal Plot SolutionKey Function & Parameters
$N < 5,000$ bivariate continuous dataStandard Scatter Plotax.scatter(x, y, alpha=0.5)
$N > 50,000$ dense continuous pointsHexagonal Binning (Hexbin)ax.hexbin(x, y, gridsize=35, cmap='Blues')
Linear trend confirmation with confidence bandSeaborn Regression Plotsns.regplot(data=df, x='x', y='y')
4 continuous variables in one viewMultivariate Bubble Chartax.scatter(x, y, s=size, c=color) + colorbar
Interactive Challenge 06

Architectural Plot Selection Drill

Select the optimal visualization strategy for each scenario. (Practice starts unselected).

Scenario 1: Evaluating 200,000 Taxi Rides for Distance vs Fare
You have 200,000 trip records. When plotted with normal dots, the center is a solid black block.
Scenario 2: Software Pricing vs Subscription Volume across 4 Competitors
Comparing price points and customer volumes for discrete distinct companies.
11

Capstone Project: SaaS CAC vs LTV Cohort Analytics

End-to-end Python pipeline with OLS regression, median crosshairs, bubble sizing, and executive export

Study this complete end-to-end Capstone script, then complete the hands-on coding drill below. Notice how the pipeline reads the raw data, fits an OLS trendline, segments into 4 quadrants, and formats the canvas.

Python • Complete Capstone Scatter Pipeline
# Pathubs Capstone: SaaS CAC vs LTV Cohort Analytics
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd

# 1. Instantiate Figure and Axes
fig, ax = plt.subplots(figsize=(11, 6.5), dpi=120)

# 2. Draw Multivariate Scatter (X, Y, Size, Color)
scatter = ax.scatter(
    df['cac'],
    df['ltv'],
    s=df['seats'] * 1.5,
    c=df['margin'],
    cmap='viridis',
    alpha=0.65,
    edgecolors='#ffffff',
    linewidths=0.6,
    label='Client Cohorts'
)

# 3. Add Colorbar
cbar = fig.colorbar(scatter, ax=ax)
cbar.set_label('Gross Margin (%)', fontsize=10, weight='bold')

# 4. Fit and Plot OLS Linear Regression Trendline
m, b = np.polyfit(df['cac'], df['ltv'], deg=1)
r = df['cac'].corr(df['ltv'])
x_line = np.linspace(df['cac'].min(), df['cac'].max(), 100)
ax.plot(x_line, m * x_line + b, color='#f43f5e', lw=2.2, ls='--', label=f'OLS Trendline (r={r:.2f})')

# 5. Add 3:1 Golden Ratio Benchmark Reference Line
ax.plot(x_line, 3.0 * x_line, color='#f59e0b', lw=1.6, ls=':', label='Golden Ratio (3:1 LTV/CAC)')

# 6. Polish Canvas
ax.set_title('SaaS CAC vs LTV Cohort Health & Unit Economics', fontsize=14, weight='bold', pad=14)
ax.set_xlabel('Customer Acquisition Cost (CAC $)', fontsize=11, weight='bold')
ax.set_ylabel('Customer Lifetime Value (LTV $)', fontsize=11, weight='bold')
ax.grid(True, linestyle='--', alpha=0.3)
ax.legend(loc='upper left', frameon=True, facecolor='#060a18', edgecolor='none')

plt.tight_layout()
plt.savefig('saas_cac_ltv_capstone.png', dpi=300, bbox_inches='tight')
plt.show()
Hands-On Coding Drill: Scatter & Trendline Construction

Write the Python code to initialize an explicit Figure/Axes, plot a scatter plot of df['spend'] vs df['revenue'] with alpha=0.5, compute Pearson correlation with .corr(), and fit an OLS regression trendline using np.polyfit().

Scatter Plot Knowledge & Mastery Assessment
Score: 0 / 8
Q1.What primary analytical relationship is a Scatter Plot specifically designed to reveal?
Q2.In Matplotlib ax.scatter(x, y, s=size), what geometric unit does the "s" parameter represent?
Q3.What is the most effective immediate visual remedy when 50,000 points create an opaque, unreadable black blob on a scatter plot (overplotting)?
Q4.What critical lesson does Anscombe's Quartet demonstrate to data analysts?
Q5.How do you calculate and plot an Ordinary Least Squares (OLS) linear trendline on a scatter plot using NumPy and Matplotlib?
Q6.What is a "high-leverage outlier" on a scatter plot?
Q7.How many variables can a well-designed 2D scatter plot effectively encode simultaneously without 3D perspective distortion?
Q8.Why is it dangerous to infer causality when a scatter plot displays a strong positive correlation (r = 0.92)?
What You Should Know Now (Competency Checklist)
Why 2D position (X, Y) offers the highest perceptual accuracy in data visualization.
How to initialize explicit Figure/Axes with fig, ax = plt.subplots().
Overcoming overplotting by tuning alpha translucency (alpha=0.2 to 0.5).
Correctly scaling bubble sizes by area in points squared ($pt^2$) rather than linear radius.
Calculating Pearson correlation df['x'].corr(df['y']) and avoiding false causality.
Fitting OLS regression lines using NumPy np.polyfit(x, y, 1).
Creating 2x2 decision quadrants with ax.axvline() and ax.axhline().
Understanding Anscombe's Quartet and why summary statistics require visual verification.