Power analysis helps students plan studies before collecting data by estimating how large a sample is needed to detect a meaningful effect. This cheat sheet connects hypothesis testing, confidence intervals, effect size, and sample size in one reference. It is useful for designing experiments, evaluating survey claims, and understanding why small studies often miss real effects.
Key Facts
- Statistical power is the probability of correctly rejecting a false null hypothesis, so .
- A Type I error occurs when a true null hypothesis is rejected, and its probability is .
- A Type II error occurs when a false null hypothesis is not rejected, and its probability is .
- For estimating a population mean with margin of error , an approximate sample size is .
- For estimating a population proportion with margin of error , an approximate sample size is .
- When is unknown for a proportion sample size calculation, use because it gives the largest required .
- Cohen's standardized mean effect size is .
- Increasing , increasing effect size, increasing , or reducing variability generally increases statistical power.
Vocabulary
- Power
- Power is the probability that a test detects a real effect when the alternative hypothesis is true.
- Significance Level
- The significance level is the probability of making a Type I error in a hypothesis test.
- Type II Error
- A Type II error happens when a test fails to reject even though is false.
- Effect Size
- Effect size measures how large a difference or relationship is in practical, often standardized, terms.
- Margin of Error
- The margin of error is the maximum expected distance between a sample estimate and the true population value at a chosen confidence level.
- Minimum Sample Size
- Minimum sample size is the smallest needed to achieve a target margin of error or power under stated assumptions.
Common Mistakes to Avoid
- Confusing and is wrong because measures false positives while measures false negatives.
- Using a smaller sample size than the calculation suggests is wrong because it can lower power and make real effects harder to detect.
- Forgetting to square the entire fraction in is wrong because sample size depends on the square of both the critical value and the margin of error.
- Using as an estimate after better prior information is available can be inefficient because the best sample size calculation should use the most reasonable expected value of .
- Treating statistical significance as practical importance is wrong because a tiny effect can be significant with a very large but still not matter in real life.
Practice Questions
- 1 A researcher wants to estimate a mean with confidence, known , and margin of error . Using , find the required sample size .
- 2 A school survey estimates a proportion with confidence and margin of error . If no prior estimate of is known, use and to find .
- 3 A test has . What is the power of the test?
- 4 Explain why increasing sample size usually increases power, even when the significance level stays the same.
Understanding Power Analysis & Sample Size Reference
Power is easiest to understand as a study's ability to separate a real pattern from ordinary random noise. Imagine testing whether a new revision routine improves test scores. Students do not all start at the same level, and scores naturally vary from day to day.
If the routine produces only a small improvement, that improvement can be hidden inside the usual spread of scores. A larger group gives a clearer picture of the average result because unusual individual scores have less influence. This is why studies of rare diseases, small learning differences, or weak product effects often need many participants.
Planning requires choices before any data are collected. Researchers first decide what difference would be important enough to detect. This is not always the smallest possible difference.
For example, a one minute reduction in a bus journey may be real but not useful, while a ten minute reduction could matter to passengers. Next, they estimate how variable the measurements are, often using earlier research or a pilot study. Greater variation means more uncertainty, so the planned sample must grow.
They then set a significance level and a target power, commonly eighty percent or ninety percent. A stricter significance level reduces the chance of a false alarm, but it usually requires more data to maintain the same ability to find a genuine effect.
Effect size helps put results on a common scale. A difference of five marks means something different in two classes whose marks have very different spreads. Standardising the difference compares it with the typical variation among students.
This makes it easier to judge whether a result is tiny, moderate, or large in context. Context matters more than a label.
A small average change in blood pressure may matter for public health when it affects millions of people. A large difference in a small classroom study may still be uncertain if the data are highly variable or the groups were not chosen fairly.
Confidence intervals give another useful view after data collection. They show a range of population values that are reasonably consistent with the sample, based on the method used. A narrow interval suggests a more precise estimate, while a wide interval shows that important uncertainty remains.
Precision improves with a larger sample, but not in a simple one for one way. To cut the margin of error in half, a study often needs about four times as many observations. Students should watch for hidden limits in sample size claims.
A huge sample cannot fix biased questions, missing groups, poor measurements, or an unfair comparison. Random sampling and random assignment solve different problems, and both should be considered when judging a study.