A normal distribution is a bell-shaped model used to describe many real data sets, such as test scores, heights, and measurement errors. This cheat sheet helps students connect the shape of the curve to mean, standard deviation, z-scores, and probability. It is especially useful for quickly estimating how unusual a value is and what percent of data falls in a given interval.
The most important ideas are that the mean is at the center, the standard deviation controls spread, and total area under the curve equals . The Empirical Rule says about of values fall within standard deviation, within , and within . A z-score, , converts any normal value into its number of standard deviations from the mean.
Percentiles and probabilities come from areas under the normal curve.
Key Facts
- A normal distribution is symmetric, bell-shaped, and centered at the mean .
- In a normal distribution, the mean, median, and mode are all equal: .
- The total area under a normal curve is , which represents of the data.
- The Empirical Rule says about of data lies between and .
- The Empirical Rule says about of data lies between and .
- The Empirical Rule says about of data lies between and .
- A z-score is calculated with and tells how many standard deviations is from the mean.
- For any normal distribution, standardizing with changes it to the standard normal distribution with and .
Vocabulary
- Normal distribution
- A symmetric bell-shaped distribution where most values are near the mean and fewer values occur farther away.
- Mean
- The center or balance point of a normal distribution, usually written as .
- Standard deviation
- A measure of spread, written as , that describes how far values typically are from the mean.
- Empirical Rule
- A rule for normal distributions stating that about , , and of data fall within , , and standard deviations of the mean.
- Z-score
- A standardized value that tells how many standard deviations a data value is above or below the mean.
- Percentile
- A location in a distribution showing the percent of data values at or below a given value.
Common Mistakes to Avoid
- Using the Empirical Rule for non-normal data is wrong because the , , and pattern only applies well to bell-shaped, approximately normal distributions.
- Forgetting that standard deviation must be positive is wrong because measures spread and cannot be less than .
- Subtracting in the wrong order for a z-score is wrong because the correct formula is , not .
- Thinking a negative z-score means an impossible value is wrong because a negative z-score only means the value is below the mean.
- Confusing area with height on the curve is wrong because probability is represented by area under the curve, not by how tall the curve is at one point.
Practice Questions
- 1 A normal distribution has mean and standard deviation . Find the interval that contains about of the data.
- 2 A test score of comes from a normal distribution with and . Calculate the z-score using .
- 3 In a normal distribution with and , estimate the percent of data between and .
- 4 Explain why two data values with the same z-score from different normal distributions have the same relative position, even if the original values are different.
Understanding Normal Distribution & Empirical Rule
Standard deviation is more than a number printed beside an average. It describes a typical distance between individual values and the center of the data. A small standard deviation means values tend to cluster tightly.
A large one means they are more spread out. Its calculation gives extra weight to large gaps because it begins by squaring each distance from the mean.
This is useful because a value far from the center should affect the measure of spread more than a value that is only slightly different. The final square root puts the answer back into the original units, such as centimeters, points, or seconds.
The Empirical Rule becomes more useful when the curve is split into smaller regions. From the mean to one standard deviation above it is about 34 percent of the data. The matching region below the mean contains another 34 percent.
Between one and two standard deviations on either side, each band contains about 13.5 percent. The narrow bands between two and three standard deviations contain about 2.35 percent each. Beyond three standard deviations, only about 0.15 percent remains in each tail.
These pieces let students estimate probabilities for intervals that do not start at the center. They can subtract areas to find the chance that a value falls between two locations.
A z-score removes the original unit and makes comparison possible. A student who scores 82 on one test cannot be compared fairly with a student who scores 82 on a different test if the tests have different averages and spreads. Their z-scores show their relative positions instead.
A positive z-score is above the mean, while a negative z-score is below it. A z-score of two means the value is unusually high compared with most observations, regardless of whether the original value was a score, height, or manufacturing measurement. Tables and calculator functions often give the area to the left of a z-score.
To find an area to the right, subtract that left area from one. To find an area between two scores, subtract the smaller left area from the larger one.
Real data do not automatically follow a normal model. A graph with strong skew, several groups, extreme outliers, or hard upper and lower limits may not fit the bell shape well. Test scores can pile up near 100 if a test is too easy.
Household income often has a long right tail because a small number of incomes are very large. In these cases, estimates from the Empirical Rule can be misleading. Before using the rule, inspect a histogram or dot plot and consider how the data were collected.
The rule describes a model, not a guarantee about every data set. Students should keep track of units, label the mean and standard deviation clearly, and state whether an answer is an estimate or a value taken from a table or calculator.