Two Categorical Variables
Section titled “Two Categorical Variables”When two categorical variables are measured on the same individuals, organize the data in a two-way table. Each cell records the count for one combination of categories.
| Category 1 | Category 2 | Total | |
|---|---|---|---|
| Group A | |||
| Group B | |||
| Total |
Marginal relative frequencies use row or column totals to describe one variable by itself. Conditional relative frequencies restrict attention to one row or column and describe the distribution of the other variable within that condition.
Simulation
Section titled “Simulation”A simulation uses random digits, technology, cards, dice, or another chance device to imitate a random process. Simulations are useful when exact probability is hard to calculate or when you want to understand long-run behavior before using formulas.
Good simulations clearly define:
- What one trial represents.
- How random outcomes are assigned to the real-world outcomes.
- What statistic is recorded after each trial.
- How many repetitions are used.
- How the simulation results answer the original question.
Sample space and experiments
Section titled “Sample space and experiments”- A random phenomenon or probability experiment is a process with outcomes that vary from trial to trial in a way that cannot be predicted with certainty in advance, but whose possible outcomes are known.
- The sample space, denoted , is the set of all possible outcomes of that experiment. An event is any subset of the sample space (a collection of one or more outcomes). The letter is standard notation; individual outcomes are often written as simple labels or ordered pairs when the experiment has multiple stages.
- A tree diagram lists stages of an experiment as branches. Multiply along a path to get the probability of that path when stages are suitably independent or conditional probabilities are marked on branches; add paths that represent the same event. Tree diagrams keep ordered outcomes visible and help avoid double-counting when the experiment is multistep.
Basic rules of probability
Section titled “Basic rules of probability”- Rule 1 (bounds): For any event ,
- Rule 2 (whole sample space): If each possible outcome is listed exactly once, the sum of their probabilities is .
- Rule 3 (impossibilility): An impossible event has probability . A sure event (the entire sample space, or any event that must happen) has probability .
- Rule 4 (odds in favor): Odds in favor of an event compare the chance occurs to the chance it does not. With complement (or ) for “not ,”
Odds are a ratio, not a probability; you can recover probabilities from odds with a little algebra when needed.
Events: complement, disjoint, union, intersection
Section titled “Events: complement, disjoint, union, intersection”- The complement of an event is the event that does not occur. It includes every outcome in that is not in . Notation includes and .
- Disjoint events (also called mutually exclusive events) cannot both occur on the same trial: they share no outcomes, so when and are disjoint.
- The union is “ or or both”—at least one of the two events happens.
- The intersection is “ and both” happen.
- Conditional probability “ given ” restricts attention to outcomes where occurred. The notation is
read as “ given .”
Two events are independent if knowing whether one occurred does not change the probability of the other (formalized below).
More probability rules
Section titled “More probability rules”Complement rule:
General addition rule (union):
If and are disjoint, then and the rule reduces to .
General multiplication rule (intersection):
From the multiplication rule you obtain the standard formula for conditional probability—the same rearrangement that describes Bayes’ theorem in tree-and-table problems:
provided .
Independence (equivalent formulations for events with positive probability):
Equivalently, independence is often written as
If that product rule fails, the events are dependent.
Random variables and probability distributions
Section titled “Random variables and probability distributions”A random variable assigns a numerical value to each outcome of a random experiment. Customarily we use capital letters such as or for the variable and lowercase for a possible value it might take.
- A discrete random variable takes a finite or countably infinite set of values (counts, “number of successes,” and so on).
- A continuous random variable takes values in an interval (time, weight, distance). Probabilities for continuous models are assigned to intervals using density and area ideas in later work; this unit focuses on the discrete case.
Discrete probability distributions
Section titled “Discrete probability distributions”A discrete probability distribution lists every possible value of together with its probability (or ). The list may appear as a table, a formula, or a probability histogram.
Let take values with probabilities . The pairing
is a valid probability distribution if
and
The first condition keeps each entry a legitimate probability; the second says exactly one of the listed values occurs (for a complete list of possibilities).
Mean of a discrete random variable
Section titled “Mean of a discrete random variable”The expected value of a discrete random variable is also called its mean and denoted when emphasis is on the distribution. It is the probability-weighted average of the possible values:
That number is a center of the distribution, but it need not be a value can actually take.
Bonus!
Section titled “Bonus!”Surprisingly, expected value is an additive property. The linearity of expectation states that:
This property is useful in games, insurance, and counting problems because it does not require the random variables to be independent. Variance rules, however, do require independence in the simple forms used in AP Statistics.
Variance of a discrete random variable
Section titled “Variance of a discrete random variable”The variance (or ) measures spread around the mean as a probability-weighted average of squared deviations:
The standard deviation is , returned to the original units of .
Combinations (binomial coefficient)
Section titled “Combinations (binomial coefficient)”A combination counts how many ways you can choose objects from distinct objects when order does not matter. The symbol is the binomial coefficient , read “ choose ”:
This expression appears in the binomial probability model (fixed independent trials, same success probability , count successes) and in many counting-based probability problems on the exam.
Binomial and geometric models
Section titled “Binomial and geometric models”A binomial random variable counts successes in a fixed number of independent trials with the same success probability. If , then
The mean and standard deviation are
A geometric random variable counts trials until the first success. If , then
Normal Distributions
Section titled “Normal Distributions”A normal distribution is a symmetric, bell-shaped distribution described by its mean and standard deviation :
To standardize a value,
Sampling Distributions and the Central Limit Theorem
Section titled “Sampling Distributions and the Central Limit Theorem”A sampling distribution is the distribution of a statistic over many random samples of the same size from the same population. It describes how a statistic varies from sample to sample.
For a statistic, pay attention to:
- Shape: normal, skewed, approximately symmetric, etc.
- Center: the mean of the statistic.
- Spread: the standard deviation of the statistic, often called the standard error.
For a sample mean,
The Central Limit Theorem says that when is large, the sampling distribution of is approximately normal, even if the population distribution is not normal, as long as observations are independent and the population is not extremely pathological.