Two Quantitative Variables
Section titled “Two Quantitative Variables”Bivariate data records two variables for every individual or object in a study. Examples include (study hours, exam score) for each student, or (latitude, January temperature) for each city.
If both variables are quantitative and the relationship looks roughly linear, we summarize direction and strength with the correlation coefficient and describe the overall trend with a least-squares regression line.
Scatterplots
Section titled “Scatterplots”A scatterplot plots each case as a point in the plane. Choose scales so that all observed - and -values fit comfortably, and label axes with variable names and units.
When you describe a scatterplot, organize your comments around three ideas: shape, direction, and strength.
Shape answers whether the overall pattern is linear (points basically follow a straight line) or nonlinear (curved, piecewise, or scattered without a simple path). Nonlinear patterns are a signal that a straight-line model may be wrong unless you transform a variable first.
Direction
Section titled “Direction”Direction describes how tends to move as increases. An upward direction is positive association and a downward direction is negative association. Clouds with no clear trend show weak or no linear association (correlation near zero is possible even when a strong nonlinear pattern exists, which is one reason you always look at the plot).
Strength
Section titled “Strength”Strength describes how tightly points follow the trend. If you imagine a line or smooth curve through the cloud, strength is about how far points deviate from that trend. Tight clouds imply strong association; wide vertical scatter implies weak association. Outliers can stand away from the bulk of the data in , in , or in both.
Pearson’s correlation coefficient
Section titled “Pearson’s correlation coefficient”Pearson’s correlation coefficient (often called the correlation is a number that measures the direction and strength of a linear relationship between two quantitative variables. It is denoted for a population and for a sample (although is almost always used). For every distribution:
The sign of matches the direction of the linear trend: for positive association, for negative association. The magnitude relates to strength for linear association only.
Formula
Section titled “Formula”For paired data with sample means and :
Intuitively, compares covariation (do and tend to be on the same side of their means together?) to how spread out and are individually.
Interpreting
Section titled “Interpreting ∣r∣”Values with mean all points fall exactly on a single straight line (perfect linear fit). As moves toward 0, the linear trend weakens.
Textbooks sometimes give rough cutoffs such as as very weak, 0.1 to 0.5 as weak-to-moderate, 0.5 to 0.85 as strong, and as very strong. Treat these as rules of thumb, not laws: context, sample size, and outliers matter. Sometimes a value of 0.4 can be classified as strong, and sometimes a value of 0.8 may be classified as weak. Correlation is not causation; confounding and lurking variables can produce strong without a direct cause-and-effect link.
Correlation is unitless and unchanged by linear rescaling (multiplying either variable by a positive constant, or adding a constant), which makes it handy for comparing relationships measured in different units.
Least-squares regression
Section titled “Least-squares regression”The linear model
Section titled “The linear model”A linear regression model describes how a response variable depends on an explanatory variable (also called a predictor or independent variable, depending on the textbook). A population-style statement often looks like
Here is the y-intercept, is the slope, and is the error term. The errors are what we hope stay small and behave reasonably once we estimate the line from data.
From a sample, we write estimated coefficients (often and , or and ) and a fitted line used for prediction. Note that is only included when you have to compare the experimental data with the model preictions.
Predicted values and residuals
Section titled “Predicted values and residuals”For a chosen , the predicted value is the height of the regression line at that . A common form is
using the least-squares estimates and from your data (notation varies).
The residual for that case is
the observed response minus the predicted response. Residuals are the data’s way of telling you where the line was too high or too low. A positive residual means the point lies above the line, and a negative residual means it lies below.
Why “least squares”?
Section titled “Why “least squares”?”The least-squares regression line is the line that minimizes the sum of squared residuals, . That criterion is mathematically tractable and can be solved without heavy use of approximations, among other mathematical advantages. In addition, it only compares the magnitudes of error, so direction does not matter.
Useful facts for AP Statistics:
- The least-squares line always passes through (), the point of means.
- The slope satisfies
where and are the sample standard deviations of and . So the sign of matches the sign of , and the steepness scales with how spread out is relative to .
Coefficient of determination
Section titled “Coefficient of determination”The coefficient of determination, , reports the fraction of the variability in that is accounted for by the linear model using . In simple linear regression with one , equals and lies between 0 and 1. Values near 1 mean the points hug the line; values near 0 mean the line explains little of how moves. is typically used instead of because it does not depend on direction, so it only shows the correlation.
High does not prove the model is appropriate (nonlinearity can still hide in residual plots), and it does not prove causation.
Technology output often includes the intercept, slope, residual standard deviation , and . In this unit, use those values descriptively: interpret the slope, make predictions when appropriate, and use residuals to judge whether a linear model is reasonable.
Influential observations and outliers
Section titled “Influential observations and outliers”An outlier in regression is often a point with an unusually large residual: the line misses it badly. An influential observation is one whose removal would substantially change the estimated slope or intercept—often a point that is extreme in (high leverage) and also off the trend. Not every outlier is influential, and not every influential point looks like a vertical outlier; inspect the plot and, when possible, recompute the line without suspect cases (sensibly and transparently).
Residual plots
Section titled “Residual plots”A residual plot graphs residuals (usually on the vertical axis) against either the predicted values or the explanatory variable . The purpose is to diagnose the fit of a linear model.
What you hope to see is a formless cloud: points scattered randomly around the horizontal axis at , with roughly constant spread across values of or .
Curved patterns mean the relationship is probably nonlinear; a linear model is a poor summary and should not be used. Fan shapes (spread grows or shrinks as changes) suggest nonconstant variance, which matters more when you move into formal inference, but is still worth mentioning when you describe real data.
Transformations to achieve linearity
Section titled “Transformations to achieve linearity”When a scatterplot shows a nonlinear trend, one strategy is to transform one or both variables so that the new relationship is more nearly linear. You then fit the line to the transformed scale and interpret conclusions in original units when you report results.
Example: if grows exponentially with , plotting against may straighten the cloud. Symbolically, if in an idealized world, then is linear in .
Common transformations
Section titled “Common transformations”- Log transformation: or for right-skewed positive responses or multiplicative growth patterns.
- Square root transformation: for count data or mild right skew where logs feel too aggressive.
- Reciprocal transformation: when larger corresponds to smaller in a rate-like way.
Always check a residual plot after transforming; the goal is a linear trend with well-behaved residuals, not a cosmetic change on the scatterplot alone.