Reading a research study is a skill for deciding whether a statistical claim is trustworthy. A study may use real data and still lead to a weak conclusion if the sample, design, or analysis is flawed. Good readers look past the headline and inspect how the evidence was collected, measured, and interpreted.
This matters in science, medicine, economics, psychology, and any field where data are used to guide decisions.
A strong reading process checks the research question, sample size, sampling method, study design, variables, possible confounders, effect size, uncertainty, and statistical significance. Randomized experiments can support stronger cause and effect claims than observational studies, but both require careful interpretation. A result can be statistically significant while still being small, biased, or not useful in practice.
The best conclusion matches the evidence without overstating what the study actually shows.
Understanding Statistics: Reading a Research Study
Start by finding the exact claim the authors set out to test. A paper may begin with a broad topic but measure only one narrow outcome. For example, a study about sleep and school performance might measure one test score after a single week.
That cannot establish every effect of sleep on learning. Check whether the researchers stated their plan before seeing the results.
A preregistered plan makes it harder to change the hypothesis or analysis after a promising pattern appears. It is useful to notice whether the conclusion refers to the same group, time period, and outcome that were actually studied.
Look closely at who entered the study and who did not. A survey of volunteers from a school website can miss students without reliable internet access or students who do not choose to respond. Even a large response count may represent a limited group.
Read the recruitment details, the eligibility rules, and the response rate. Missing data need attention too.
If many participants leave before the final measurement, the remaining group may differ in an important way. Researchers can sometimes adjust for missing data, but an adjustment depends on assumptions that may not hold.
The way a variable is measured can change the result. Self-reported exercise, screen time, pain, or income may contain memory errors or pressure to give a socially acceptable answer. A measurement tool should be reliable, meaning it gives similar results under similar conditions.
It should be valid, meaning it measures the intended idea. In an experiment, participants may behave differently if they know which treatment they received.
Researchers use placebo groups and blinding to reduce this problem. Blinding means that participants, researchers, or both do not know who received the treatment during key parts of the study.
Numbers need context before they support a conclusion. A small p-value shows that the data would be unusual under a particular no-effect model. It does not show that the finding is important, true, or likely to appear again.
Examine the confidence interval because it shows how precise the estimate is. A wide interval means several effect sizes remain plausible. Check how many comparisons were made.
If researchers test many outcomes or many groups, some apparently significant results can occur by chance. Good studies report all planned analyses, not only the most striking result.
Finally, consider whether the result applies beyond the study. A medicine tested in healthy adults may work differently for children, older people, or people with other health conditions. A classroom method tested by one skilled teacher may not produce the same result elsewhere.
Replication matters because a second study can reveal whether a finding survives new participants, new researchers, and slightly different conditions. When reading a headline or sharing a study online, use language that matches the evidence. Say that a study found an association when it did not establish cause.
Say that results are preliminary when uncertainty remains. Careful wording is part of careful statistics.
Key Facts
- Sample size matters because larger samples usually reduce random sampling error, but they do not fix biased sampling.
- Margin of error often decreases like 1/sqrt(n), where n is the sample size.
- A p-value is P(data this extreme or more extreme | null hypothesis is true), not the probability that the null hypothesis is true.
- A 95% confidence interval gives a range of plausible values for a population parameter based on the study method.
- Effect size measures how large a difference or relationship is, such as difference in means = mean treatment - mean control.
- Correlation does not prove causation because confounders, reverse causation, or selection effects may explain the pattern.
Vocabulary
- Sample
- A sample is the group of individuals, objects, or cases actually measured in a study.
- Population
- A population is the larger group that the researchers want to draw conclusions about.
- Confounder
- A confounder is a variable related to both the explanatory variable and the outcome that can distort the apparent relationship.
- Statistical significance
- Statistical significance means the observed result would be unlikely under a specified null hypothesis, often judged using a p-value cutoff.
- Effect size
- Effect size is a measure of the magnitude of a difference, association, or treatment impact.
Common Mistakes to Avoid
- Treating a small sample as automatically convincing is wrong because random variation can strongly affect small studies and make results unstable.
- Reading a p-value as the chance the claim is true is wrong because a p-value is calculated assuming the null hypothesis, not the study claim.
- Ignoring the study design is wrong because an observational study usually cannot establish causation without strong assumptions and controls.
- Focusing only on statistical significance is wrong because a tiny effect can be significant in a large sample but still be unimportant in practice.
Practice Questions
- 1 A survey reports that 54 out of 300 students prefer online homework. What sample proportion prefers online homework?
- 2 A treatment group has a mean score of 82 and a control group has a mean score of 76. What is the difference in means, and which group scored higher?
- 3 A study finds that people who drink more coffee have lower rates of a disease. Explain why this result alone does not prove that coffee prevents the disease.