Probability distributions describe how likely different outcomes are for a random variable. This cheat sheet helps students compare discrete and continuous models, choose the right formula, and interpret probabilities from tables or graphs. It is useful for homework, tests, and data investigations where probabilities must be calculated and explained.
The core ideas are that probabilities must total for discrete distributions and total area must equal for continuous distributions. Expected value gives the long-run average, while variance and standard deviation measure spread. Common models such as the binomial and normal distributions use parameters like , , , and to describe shape, center, and variability.
Key Facts
- For any probability distribution, and the probabilities of all mutually exclusive outcomes in the sample space add to .
- For a discrete random variable, the mean or expected value is .
- For a discrete random variable, the variance is and the standard deviation is .
- For a binomial random variable with trials and success probability , for .
- A binomial distribution has mean and standard deviation .
- For a continuous random variable, equals the area under the density curve from to , and .
- A normal random variable is standardized by , which converts into the number of standard deviations from the mean.
- If , then about , , and of values lie within , , and standard deviations of .
Vocabulary
- Probability distribution
- A probability distribution is a model that assigns probabilities to the possible values of a random variable so the total probability is .
- Random variable
- A random variable is a variable, usually written as , whose value depends on the outcome of a chance process.
- Probability mass function
- A probability mass function gives for each possible value of a discrete random variable.
- Probability density function
- A probability density function is a curve for a continuous random variable where probabilities are found as areas under the curve.
- Expected value
- Expected value is the long-run average outcome of a random variable, calculated for discrete data by .
- Standard deviation
- Standard deviation is a measure of typical distance from the mean, calculated as .
Common Mistakes to Avoid
- Adding probabilities to more than is wrong because a valid probability distribution must have total probability equal to .
- Using the binomial formula when trials are not independent is wrong because assumes the same on every trial.
- Treating as an area for a continuous distribution is wrong because a single point has no width, so .
- Confusing variance and standard deviation is wrong because variance is while standard deviation is .
- Using like a z-score is wrong because standard normal probabilities require before using a normal table or calculator.
Practice Questions
- 1 A discrete random variable has values , , and with probabilities , , and . Find .
- 2 Let . Find using .
- 3 A test score is normally distributed with and . Find the z-score for and interpret its meaning.
- 4 Explain whether the number of students absent from a class on a school day is better modeled as a discrete or continuous random variable, and justify your choice.
Understanding Probability Distributions
A distribution is a model, not a guarantee about one result. It describes a pattern that would appear after many repetitions or across a large group. The first job is to define the random variable clearly.
For example, a variable might count the number of defective items in a box, record a travel time, or measure a student height. Counts are usually discrete because only whole-number results make sense.
Measurements are usually continuous because values can fall anywhere within a range. This distinction affects the graph, the calculation method, and the meaning of a single value.
The expected value is useful because it gives a fair long-run outcome, even when that outcome cannot actually occur. A game can have an expected winning of two dollars even if every possible prize is a whole number of dollars. This helps students judge lotteries, insurance costs, and repeated business decisions.
Variance and standard deviation describe how dependable that average is. Two situations can have the same expected value but very different risk. One may produce results close to the average most of the time.
Another may produce extreme gains or losses. Standard deviation uses the same units as the original data, so it is often easier to interpret than variance.
A binomial model fits only under specific conditions. There must be a fixed number of trials. Each trial must have two categories, often called success and failure.
The chance of success must stay the same, and trials should be independent. Tossing a coin several times is a good example. Drawing cards without replacement usually is not, because each draw changes the deck.
In real surveys, independence can fail when classmates influence one another or when people from the same household are selected. The shape of a binomial distribution depends on the chance of success.
When success is rare, the graph tends to bunch near zero with a longer tail toward larger counts. When success is about equally likely as failure, the graph is more balanced.
Normal distributions are especially useful for measurements influenced by many small effects. Test scores, manufacturing sizes, and biological measurements can sometimes be close to normal, but real data should be checked rather than assumed to be normal. A histogram or plot can reveal skewness, gaps, or unusual outliers.
A z-score puts values from different scales onto a common scale. This makes it possible to compare a score on one test with a score on another test, even when their averages and spreads differ. Positive z-scores are above the mean, while negative z-scores are below it.
When using a normal table or calculator, pay close attention to whether it gives area to the left, area to the right, or area between two values. Many errors come from finding the correct area but reporting the wrong probability.