Skip to content

Unit 2: Probability, Random Variables, and Probability Distributions

AP Stats cheatsheet

Open ↗

Loading…

When two categorical variables are measured on the same individuals, organize the data in a two-way table. Each cell records the count for one combination of categories.

Category 1Category 2Total
Group Aaabba+ba+b
Group Bccddc+dc+d
Totala+ca+cb+db+dnn

Marginal relative frequencies use row or column totals to describe one variable by itself. Conditional relative frequencies restrict attention to one row or column and describe the distribution of the other variable within that condition.


A simulation uses random digits, technology, cards, dice, or another chance device to imitate a random process. Simulations are useful when exact probability is hard to calculate or when you want to understand long-run behavior before using formulas.

Good simulations clearly define:

  1. What one trial represents.
  2. How random outcomes are assigned to the real-world outcomes.
  3. What statistic is recorded after each trial.
  4. How many repetitions are used.
  5. How the simulation results answer the original question.

  • A random phenomenon or probability experiment is a process with outcomes that vary from trial to trial in a way that cannot be predicted with certainty in advance, but whose possible outcomes are known.
  • The sample space, denoted SS, is the set of all possible outcomes of that experiment. An event is any subset of the sample space (a collection of one or more outcomes). The letter SS is standard notation; individual outcomes are often written as simple labels or ordered pairs when the experiment has multiple stages.
  • A tree diagram lists stages of an experiment as branches. Multiply along a path to get the probability of that path when stages are suitably independent or conditional probabilities are marked on branches; add paths that represent the same event. Tree diagrams keep ordered outcomes visible and help avoid double-counting when the experiment is multistep.
startABA\CA\CcB\CB\CcP(A)P(B)P(CjA)

  • Rule 1 (bounds): For any event AA,
0≤P(A)≤10 \le P(A) \le 1
  • Rule 2 (whole sample space): If each possible outcome is listed exactly once, the sum of their probabilities is 11.
  • Rule 3 (impossibilility): An impossible event has probability 00. A sure event (the entire sample space, or any event that must happen) has probability 11.
  • Rule 4 (odds in favor): Odds in favor of an event AA compare the chance AA occurs to the chance it does not. With complement A′A' (or ACA^C) for “not AA,”
Odds in favor of A=P(A)P(A′)\text{Odds in favor of } A = \frac{P(A)}{P(A')}

Odds are a ratio, not a probability; you can recover probabilities from odds with a little algebra when needed.


Events: complement, disjoint, union, intersection

Section titled “Events: complement, disjoint, union, intersection”
  • The complement of an event AA is the event that AA does not occur. It includes every outcome in SS that is not in AA. Notation includes A′A' and ACA^C.
  • Disjoint events (also called mutually exclusive events) cannot both occur on the same trial: they share no outcomes, so P(A∩B)=0P(A \cap B) = 0 when AA and BB are disjoint.
  • The union A∪BA \cup B is “AA or BB or both”—at least one of the two events happens.
  • The intersection A∩BA \cap B is “AA and BB both” happen.
  • Conditional probability “AA given BB” restricts attention to outcomes where BB occurred. The notation is
A∣BA \mid B

read as “AA given BB.”

Two events are independent if knowing whether one occurred does not change the probability of the other (formalized below).


Complement rule:

P(A′)=1−P(A)P(A') = 1 - P(A)

General addition rule (union):

P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B)

If AA and BB are disjoint, then P(A∩B)=0P(A \cap B) = 0 and the rule reduces to P(A∪B)=P(A)+P(B)P(A \cup B) = P(A) + P(B).

General multiplication rule (intersection):

P(A∩B)=P(A) P(B∣A)=P(B) P(A∣B)P(A \cap B) = P(A)\,P(B \mid A) = P(B)\,P(A \mid B)

From the multiplication rule you obtain the standard formula for conditional probability—the same rearrangement that describes Bayes’ theorem in tree-and-table problems:

P(A∣B)=P(A∩B)P(B)P(A \mid B) = \frac{P(A \cap B)}{P(B)}

provided P(B)>0P(B) > 0.

Independence (equivalent formulations for events with positive probability):

P(A∣B)=P(A)andP(B∣A)=P(B)P(A \mid B) = P(A) \quad \text{and} \quad P(B \mid A) = P(B)

Equivalently, independence is often written as

P(A∩B)=P(A) P(B)P(A \cap B) = P(A)\,P(B)

If that product rule fails, the events are dependent.


Random variables and probability distributions

Section titled “Random variables and probability distributions”

A random variable assigns a numerical value to each outcome of a random experiment. Customarily we use capital letters such as XX or YY for the variable and lowercase xx for a possible value it might take.

  • A discrete random variable takes a finite or countably infinite set of values (counts, “number of successes,” and so on).
  • A continuous random variable takes values in an interval (time, weight, distance). Probabilities for continuous models are assigned to intervals using density and area ideas in later work; this unit focuses on the discrete case.

A discrete probability distribution lists every possible value xix_i of XX together with its probability P(xi)P(x_i) (or P(X=xi)P(X = x_i)). The list may appear as a table, a formula, or a probability histogram.

Let XX take values x1,x2,…,xnx_1, x_2, \ldots, x_n with probabilities P(x1),P(x2),…,P(xn)P(x_1), P(x_2), \ldots, P(x_n). The pairing

{(x1,P(x1)),(x2,P(x2)),…,(xn,P(xn))}\{(x_1, P(x_1)), (x_2, P(x_2)), \ldots, (x_n, P(x_n))\}

is a valid probability distribution if

0≤P(xi)≤1for all i=1,2,…,n0 \le P(x_i) \le 1 \quad \text{for all } i = 1, 2, \ldots, n

and

∑i=1nP(xi)=1\sum_{i=1}^{n} P(x_i) = 1

The first condition keeps each entry a legitimate probability; the second says exactly one of the listed values occurs (for a complete list of possibilities).


The expected value E(X)E(X) of a discrete random variable XX is also called its mean and denoted μX\mu_X when emphasis is on the distribution. It is the probability-weighted average of the possible values:

μ=E(X)=∑i=1nxi P(xi)\mu = E(X) = \sum_{i=1}^{n} x_i\, P(x_i)

That number is a center of the distribution, but it need not be a value XX can actually take.

Surprisingly, expected value is an additive property. The linearity of expectation states that:

E(X1+X2+...+Xn)=E(X1)+E(X2)+...+E(Xn)E(X_1 + X_2 + ... + X_n) = E(X_1) + E(X_2) + ... + E(X_n)

This property is useful in games, insurance, and counting problems because it does not require the random variables to be independent. Variance rules, however, do require independence in the simple forms used in AP Statistics.


The variance σ2\sigma^2 (or Var⁡(X)\operatorname{Var}(X)) measures spread around the mean as a probability-weighted average of squared deviations:

σ2=∑i=1n(xi−μ)2 P(xi)\sigma^2 = \sum_{i=1}^{n} (x_i - \mu)^2\, P(x_i)

The standard deviation is σ=σ2\sigma = \sqrt{\sigma^2}, returned to the original units of XX.


A combination counts how many ways you can choose rr objects from nn distinct objects when order does not matter. The symbol is the binomial coefficient (nr)\binom{n}{r}, read “nn choose rr”:

(nr)=n!r! (n−r)!\binom{n}{r} = \frac{n!}{r!\,(n-r)!}

This expression appears in the binomial probability model (fixed nn independent trials, same success probability pp, count successes) and in many counting-based probability problems on the exam.


A binomial random variable counts successes in a fixed number of independent trials with the same success probability. If X∼Binomial⁡(n,p)X \sim \operatorname{Binomial}(n,p), then

P(X=k)=(nk)pk(1−p)n−k.P(X=k)=\binom{n}{k}p^k(1-p)^{n-k}.

The mean and standard deviation are

μX=np,σX=np(1−p).\mu_X=np,\qquad \sigma_X=\sqrt{np(1-p)}.

A geometric random variable counts trials until the first success. If X∼Geometric⁡(p)X \sim \operatorname{Geometric}(p), then

P(X=k)=(1−p)k−1p.P(X=k)=(1-p)^{k-1}p.
1234567binomialgeometriccountprobability

A normal distribution is a symmetric, bell-shaped distribution described by its mean μ\mu and standard deviation σ\sigma:

X∼N(μ,σ).X \sim N(\mu,\sigma).

To standardize a value,

z=x−μσ.z=\frac{x-\mu}{\sigma}.
¡3¡2¡10123normalmodelzdensity

Sampling Distributions and the Central Limit Theorem

Section titled “Sampling Distributions and the Central Limit Theorem”

A sampling distribution is the distribution of a statistic over many random samples of the same size from the same population. It describes how a statistic varies from sample to sample.

For a statistic, pay attention to:

  • Shape: normal, skewed, approximately symmetric, etc.
  • Center: the mean of the statistic.
  • Spread: the standard deviation of the statistic, often called the standard error.

For a sample mean,

μxˉ=μ,σxˉ=σn.\mu_{\bar{x}}=\mu,\qquad \sigma_{\bar{x}}=\frac{\sigma}{\sqrt{n}}.

The Central Limit Theorem says that when nn is large, the sampling distribution of xˉ\bar{x} is approximately normal, even if the population distribution is not normal, as long as observations are independent and the population is not extremely pathological.