Skip to content

Unit 5: Regression Analysis

AP Stats cheatsheet

Open ↗

Loading…

Bivariate data records two variables for every individual or object in a study. Examples include (study hours, exam score) for each student, or (latitude, January temperature) for each city.

If both variables are quantitative and the relationship looks roughly linear, we summarize direction and strength with the correlation coefficient and describe the overall trend with a least-squares regression line.


A scatterplot plots each case as a point (x,y)(x, y) in the plane. Choose scales so that all observed xx- and yy-values fit comfortably, and label axes with variable names and units.

positiveassociationnegativeassociationxy

When you describe a scatterplot, organize your comments around three ideas: shape, direction, and strength.

Shape answers whether the overall pattern is linear (points basically follow a straight line) or nonlinear (curved, piecewise, or scattered without a simple path). Nonlinear patterns are a signal that a straight-line model may be wrong unless you transform a variable first.

Direction describes how yy tends to move as xx increases. An upward direction is positive association and a downward direction is negative association. Clouds with no clear trend show weak or no linear association (correlation near zero is possible even when a strong nonlinear pattern exists, which is one reason you always look at the plot).

Strength describes how tightly points follow the trend. If you imagine a line or smooth curve through the cloud, strength is about how far points deviate from that trend. Tight clouds imply strong association; wide vertical scatter implies weak association. Outliers can stand away from the bulk of the data in xx, in yy, or in both.


Pearson’s correlation coefficient (often called the correlation is a number that measures the direction and strength of a linear relationship between two quantitative variables. It is denoted ρ\rho for a population and rr for a sample (although rr is almost always used). For every distribution:

−1≤r≤1-1 \le r \le 1

The sign of rr matches the direction of the linear trend: r>0r > 0 for positive association, r<0r < 0 for negative association. The magnitude ∣r∣\lvert r\rvert relates to strength for linear association only.

For paired data (xi,yi)(x_i, y_i) with sample means xˉ\bar{x} and yˉ\bar{y}:

r=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2∑(yi−yˉ)2r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}

Intuitively, rr compares covariation (do xx and yy tend to be on the same side of their means together?) to how spread out xx and yy are individually.

Values with ∣r∣=1\lvert r\rvert = 1 mean all points fall exactly on a single straight line (perfect linear fit). As ∣r∣\lvert r\rvert moves toward 0, the linear trend weakens.

Textbooks sometimes give rough cutoffs such as ∣r∣<0.1\lvert r\rvert < 0.1 as very weak, 0.1 to 0.5 as weak-to-moderate, 0.5 to 0.85 as strong, and ∣r∣>0.85\lvert r\rvert > 0.85 as very strong. Treat these as rules of thumb, not laws: context, sample size, and outliers matter. Sometimes a ∣r∣\lvert r\rvert value of 0.4 can be classified as strong, and sometimes a ∣r∣\lvert r\rvert value of 0.8 may be classified as weak. Correlation is not causation; confounding and lurking variables can produce strong ∣r∣\lvert r\rvert without a direct cause-and-effect link.

Correlation is unitless and unchanged by linear rescaling (multiplying either variable by a positive constant, or adding a constant), which makes it handy for comparing relationships measured in different units.


A linear regression model describes how a response variable YY depends on an explanatory variable XX (also called a predictor or independent variable, depending on the textbook). A population-style statement often looks like

Y=β0+β1X+ϵY = \beta_0 + \beta_1 X + \epsilon

Here β0\beta_0 is the y-intercept, β1\beta_1 is the slope, and ϵ\epsilon is the error term. The errors ϵ\epsilon are what we hope stay small and behave reasonably once we estimate the line from data.

From a sample, we write estimated coefficients (often b0b_0 and b1b_1, or β^0\hat{\beta}_0 and β^1\hat{\beta}_1) and a fitted line used for prediction. Note that ϵ\epsilon is only included when you have to compare the experimental data with the model preictions.

For a chosen xx, the predicted value y^\hat{y} is the height of the regression line at that xx. A common form is

y^=b0+b1x\hat{y} = b_0 + b_1 x

using the least-squares estimates b0b_0 and b1b_1 from your data (notation varies).

The residual for that case is

ϵ=y−y^\epsilon = y - \hat{y}

the observed response minus the predicted response. Residuals are the data’s way of telling you where the line was too high or too low. A positive residual means the point lies above the line, and a negative residual means it lies below.

The least-squares regression line is the line that minimizes the sum of squared residuals, ∑ϵi2\sum \epsilon_i^2. That criterion is mathematically tractable and can be solved without heavy use of approximations, among other mathematical advantages. In addition, it only compares the magnitudes of error, so direction does not matter.

Useful facts for AP Statistics:

  • The least-squares line always passes through (xˉ,yˉ\bar{x}, \bar{y}), the point of means.
  • The slope satisfies
b1=r(sysx)b_1 = r \left( \frac{s_y}{s_x} \right)

where sxs_x and sys_y are the sample standard deviations of xx and yy. So the sign of b1b_1 matches the sign of rr, and the steepness scales with how spread out yy is relative to xx.

The coefficient of determination, R2R^2, reports the fraction of the variability in yy that is accounted for by the linear model using xx. In simple linear regression with one xx, R2R^2 equals r2r^2 and lies between 0 and 1. Values near 1 mean the points hug the line; values near 0 mean the line explains little of how yy moves. R2R^2 is typically used instead of rr because it does not depend on direction, so it only shows the correlation.

High R2R^2 does not prove the model is appropriate (nonlinearity can still hide in residual plots), and it does not prove causation.

Technology output often includes the intercept, slope, residual standard deviation ss, and R2R^2. In this unit, use those values descriptively: interpret the slope, make predictions when appropriate, and use residuals to judge whether a linear model is reasonable.

An outlier in regression is often a point with an unusually large residual: the line misses it badly. An influential observation is one whose removal would substantially change the estimated slope or intercept—often a point that is extreme in xx (high leverage) and also off the trend. Not every outlier is influential, and not every influential point looks like a vertical outlier; inspect the plot and, when possible, recompute the line without suspect cases (sensibly and transparently).


A residual plot graphs residuals (usually on the vertical axis) against either the predicted values y^\hat{y} or the explanatory variable xx. The purpose is to diagnose the fit of a linear model.

randomscattermeansalinearmodelisreasonable¯ttedvalueresidual

What you hope to see is a formless cloud: points scattered randomly around the horizontal axis at ϵ=0\epsilon = 0, with roughly constant spread across values of xx or y^\hat{y}.

Curved patterns mean the relationship is probably nonlinear; a linear model is a poor summary and should not be used. Fan shapes (spread grows or shrinks as xx changes) suggest nonconstant variance, which matters more when you move into formal inference, but is still worth mentioning when you describe real data.


When a scatterplot shows a nonlinear trend, one strategy is to transform one or both variables so that the new relationship is more nearly linear. You then fit the line to the transformed scale and interpret conclusions in original units when you report results.

Example: if yy grows exponentially with xx, plotting ln⁡(y)\ln(y) against xx may straighten the cloud. Symbolically, if y=aekxy = a e^{kx} in an idealized world, then ln(y)=ln(a)+kxln(y) = ln(a) + kx is linear in xx.

  • Log transformation: z=ln(y)z = ln(y) or z=log10(y)z = log_{10}(y) for right-skewed positive responses or multiplicative growth patterns.
  • Square root transformation: z=yz = \sqrt{y} for count data or mild right skew where logs feel too aggressive.
  • Reciprocal transformation: z=1/yz = 1/y when larger xx corresponds to smaller yy in a rate-like way.

Always check a residual plot after transforming; the goal is a linear trend with well-behaved residuals, not a cosmetic change on the scatterplot alone.