Correlation means two variables tend to move together, while causation means a change in one variable directly produces a change in another. This distinction matters because data can show strong patterns that do not prove one thing caused the other. In science, medicine, economics, and daily decision making, confusing correlation with causation can lead to wrong conclusions.
A graph can suggest a relationship, but it cannot by itself prove the reason for that relationship.
A correlation may happen because A causes B, B causes A, a third variable affects both, or the pattern is a coincidence. Lurking and confounding variables are especially important because they can create or hide relationships in observed data. Experiments help establish causation by randomly assigning treatments and controlling other factors.
Good statistical reasoning asks not only whether two variables are related, but why the relationship exists.
Understanding Statistics: Correlation vs Causation
A useful way to think about a causal claim is to imagine an intervention. If a school changes one feature, such as the amount of homework, what would happen if everything else stayed as similar as possible. This is different from merely noticing that students with more homework have higher scores.
Those students may attend different classes, have more support at home, or already be strong readers. Time order matters too.
A supposed cause must happen before its effect. Even then, earlier events can be signs of an underlying process rather than the true source of a later change.
Real data often come from groups that differ before any measurement begins. People who choose to exercise may differ from non-exercisers in income, health, diet, sleep, and access to medical care. Comparing the two groups can make exercise look more or less powerful than it really is.
This is called selection bias. Another trap appears when data are combined across groups.
A trend in the full data set can weaken, disappear, or reverse after separating students by grade level, neighborhood, or prior achievement. Looking at the groups inside the total is often necessary before making a conclusion.
Well-designed experiments try to create a fair comparison. Researchers give one group a treatment and another group a comparison condition. Random assignment makes it less likely that one group starts out healthier, wealthier, or more motivated by chance.
Placebos can prevent expectations from changing results. Blinding can keep participants or researchers from knowing who received which condition while outcomes are measured. These details matter because human expectations can affect reports of pain, effort, symptoms, and performance.
Some questions cannot be tested this way because it would be unsafe or unfair to assign harmful conditions. In those cases, researchers use long-term studies, matched groups, natural events, and evidence from several independent studies.
When reading a graph or headline, first identify who was measured and how they were chosen. Check whether the relationship is based on a small sample, a short time period, or self-reported answers. Look for outliers, curved patterns, and separate clusters, since one summary number can hide these features.
A strong association can still have a small practical effect, while a weak association may matter greatly when many people are affected. Pay attention to the wording of claims. Phrases such as linked to or associated with fit observational evidence.
Words such as causes require much stronger support. Good statistical reasoning keeps the pattern, the possible mechanism, and the quality of the evidence separate.
Key Facts
- Correlation measures association, not proof of cause and effect.
- Correlation coefficient r ranges from -1 to 1.
- r = 1 means perfect positive linear correlation, r = -1 means perfect negative linear correlation, and r = 0 means no linear correlation.
- Causation means changing X produces a change in Y, often written as X causes Y.
- A confounding variable is related to both X and Y and can make a causal claim misleading.
- Randomized experiments are the strongest common method for testing causation because random assignment helps balance other variables.
Vocabulary
- Correlation
- A statistical relationship in which two variables tend to change together in a consistent pattern.
- Causation
- A cause and effect relationship in which a change in one variable directly produces a change in another variable.
- Confounding variable
- A third variable that is connected to both the possible cause and the outcome, making the relationship harder to interpret.
- Lurking variable
- An unmeasured variable that influences the variables being studied and may explain an observed pattern.
- Reverse causation
- A situation in which the outcome may actually be causing the supposed cause rather than the other way around.
Common Mistakes to Avoid
- Assuming a strong correlation proves causation is wrong because a high r value only shows a pattern between variables, not the mechanism behind it.
- Ignoring a possible confounding variable is wrong because a third factor may be responsible for changes in both variables.
- Treating time order as complete proof is wrong because even if X happens before Y, other explanations may still account for the result.
- Using observational data as if it were a randomized experiment is wrong because observed groups may differ in important ways before any treatment or exposure occurs.
Practice Questions
- 1 A study of 12 cities finds that ice cream sales and drowning incidents have correlation coefficient r = 0.86. Does this number prove that ice cream sales cause drowning incidents? Identify one likely confounding variable.
- 2 A researcher records hours studied and exam score for 20 students and finds r = 0.72. If one student studied 2 hours and scored 68, while another studied 6 hours and scored 88, calculate the change in score per extra study hour between these two students.
- 3 A school finds that students who attend a tutoring program earn higher math grades than students who do not. Explain why this observation alone does not prove the tutoring caused the higher grades, and describe an experiment that would better test causation.