The Central Limit Theorem explains why normal curves appear so often in statistics, even when the original population is not normal. This cheat sheet helps students recognize when sample means or sample proportions can be modeled with an approximately normal distribution. It is useful for solving probability problems, checking conditions, and preparing for inference topics like confidence intervals and hypothesis tests.
Key Facts
- For sample means, the Central Limit Theorem says that if is large enough, the sampling distribution of is approximately normal with mean and standard deviation .
- The standard error of the sample mean is , which measures the typical distance between and .
- A common rule of thumb is that the sampling distribution of is approximately normal when , especially if the population is not strongly skewed.
- If the original population is normal, then the sampling distribution of is normal for any sample size .
- For sample proportions, the sampling distribution of is approximately normal with mean and standard error when and .
- To standardize a sample mean, use when the population standard deviation is known.
- Larger sample sizes make the standard error smaller because decreases as increases.
- The Central Limit Theorem describes the distribution of sample statistics, not the shape of the original population data.
Vocabulary
- Central Limit Theorem
- A theorem stating that the sampling distribution of a mean or proportion becomes approximately normal as the sample size becomes large enough.
- Sampling Distribution
- The probability distribution of a statistic, such as or , calculated from many samples of the same size.
- Sample Mean
- The average value from a sample, written as .
- Standard Error
- The standard deviation of a sampling distribution, such as for sample means.
- Sample Proportion
- The fraction of a sample with a certain characteristic, written as .
- Normal Approximation
- The use of a normal distribution to estimate probabilities for a sampling distribution when the required conditions are met.
Common Mistakes to Avoid
- Confusing the population distribution with the sampling distribution is wrong because the Central Limit Theorem describes the behavior of statistics like , not the original data values.
- Using instead of for sample mean problems is wrong because sample means vary less than individual observations.
- Assuming always guarantees accuracy is wrong because strong skewness or extreme outliers may require a larger sample size.
- Forgetting the success-failure condition for proportions is wrong because is not safely normal unless and .
- Thinking a larger sample size changes the mean of the sampling distribution is wrong because stays the same while gets smaller.
Practice Questions
- 1 A population has mean and standard deviation . For samples of size , find and .
- 2 A population has and sample size . Check whether the normal approximation for is reasonable, then find .
- 3 A population has mean and standard deviation . For , calculate the z-score for a sample mean of using .
- 4 Explain why increasing the sample size makes the sampling distribution of narrower but does not change its center.
Understanding Central Limit Theorem Reference
A sampling distribution is built by imagining the same sampling process many times. Start with one population, such as the weights of every apple delivered to a store. Take a random sample, find its average weight, then put those apples back in the population in the thought experiment.
Repeat with many new random samples of the same size. The list of sample averages has its own pattern. This pattern is not a graph of individual apple weights.
It shows how much an average changes from sample to sample. Averaging balances out many random high and low values, which is why sample averages are usually less spread out than individual measurements.
Randomness and independence are essential. A large sample cannot fix a biased method. For example, a survey of students leaving a school sports event may overrepresent students who enjoy sports.
Its sample proportion may be precise, meaning it changes little from one similar survey to another, yet still be inaccurate for the whole school. Students should check how the sample was selected before using any normal model.
When sampling without replacement from a finite group, a common check is that the sample should be no more than one tenth of the population. This helps ensure that choosing one person does not greatly change the chance of choosing the next person.
The shape of the original data still matters for small samples. A few extreme values can pull an average far from the middle. Income, waiting times, and online video views often have long right tails.
In these cases, a sample of only a few observations may produce averages with an uneven pattern. Looking at a histogram, dot plot, or box plot gives useful evidence. Strong skewness, separate clusters, and outliers are warning signs.
Larger samples reduce the influence of one unusual observation, but they do not remove bias or errors in the data. For proportions, the success and failure counts matter because a normal curve is a poor model when one outcome is very rare in the sample.
A z score tells how far a sample result is from its expected center in units of standard error. This changes a difference into a common scale. A sample average that is two standard errors above the population mean is fairly unusual under a sound random process.
Tables, calculators, or software can then turn that position into a probability. This reasoning appears in election polls, quality control, medical studies, and classroom experiments. In many real studies the population standard deviation is unknown.
Students then estimate it from the sample and later learn methods based on the t distribution. The main habit is to separate an individual value from a sample statistic, then check the conditions before trusting a probability calculation.