THE SHORT ANSWER
Correlation describes how measured variables vary together. A causal claim says that changing one factor would change an outcome under specified conditions. A pattern can motivate that claim, but study design and alternative explanations determine how well it is supported.
What to remember
- A strong relationship can come from a common cause, selection, reverse direction, or other factors.
- Ask how the comparison groups were created before interpreting a numerical result as an effect.
- Correlation can still support prediction and investigation even when it does not identify a causal mechanism.
Side by side
| Question | Correlation or association | Causation |
|---|---|---|
| Main question | Do the measured variables vary together? | Would changing this factor change the outcome? |
| Direction | A numerical correlation does not identify causal direction | Specifies a direction from cause to outcome |
| Evidence needed | Suitable observations and an appropriate measure | A defensible design and assumptions addressing alternative explanations |
| Careful wording | Was associated with a difference | Changed the outcome, within the study's supported scope |
Separate an observed pattern from an intervention
Suppose students who attend more optional tutoring sessions tend to earn higher final scores. That is an association between attendance and performance. It does not yet answer what would happen if a particular student were assigned extra tutoring. Students chose whether to attend, and the reasons behind that choice may also relate to their scores.
A causal question concerns a change: for comparable students in specified conditions, what difference would receiving tutoring make? We cannot observe the same person both receiving and not receiving the same intervention at the same time. Research designs build a credible comparison to address that missing alternative, rather than treating every observed difference as an effect.
An invented tutoring result with a hidden difference
Consider a fictional class with 40 students who have strong prior preparation and 40 with limited prior preparation. Within the strongly prepared group, 30 attend tutoring and 10 do not; both subgroups average 85 on the final. Within the less prepared group, 10 attend and 30 do not; both subgroups average 65. These values are deliberately constructed to isolate a comparison problem.
Across all attendees, the mean is (30 × 85 + 10 × 65) ÷ 40 = 80. Across all non-attendees, it is (10 × 85 + 30 × 65) ÷ 40 = 70. The pooled data show a ten-point advantage for tutoring attendance. Yet the difference is zero within each preparation group in this example.
The groups differ in their composition: prepared students make up three quarters of attendees but only one quarter of non-attendees. Prior preparation is a plausible confounding factor because it relates to both attendance and scores. The arithmetic does not prove that real tutoring is ineffective. It proves that this particular pooled difference can exist without a within-group difference.
Even comparing students within preparation groups would not settle every real causal question. Motivation, available time, course choices, and the quality of tutoring could still differ. Statistical adjustment helps only to the extent that the relevant variables, measurements, and assumptions are appropriate.
The class, scores, and attendance counts are fictional. They illustrate confounding and are not a claim about any tutoring program.
Try at least three explanations before choosing one
First, consider a common cause: an earlier factor may influence both variables. Second, consider reverse direction: a developing outcome may influence the supposed cause. Students struggling with a subject may seek more tutoring, which could make helpful support appear associated with lower scores. Third, consider selection: the people measured may differ from the population the headline describes.
Measurement can also create misleading patterns. If attendance is measured accurately but prior preparation is reduced to one noisy question, adjusting for preparation may leave substantial differences between the groups. And if many relationships were searched, a striking result may have been selected from a much larger set of unremarkable comparisons.
- What happened first, and was that timing actually recorded?
- What could influence both the proposed cause and the outcome?
- Who entered or left the sample, and why?
- Were the variables and comparison selected before seeing the results?
Understand what random assignment contributes
Penn State's introductory statistics materials distinguish observational studies from randomized experiments. Random assignment creates treatment groups through chance rather than participant choice, helping address systematic differences in their starting characteristics. It is not the same as randomly sampling people from a population: sampling concerns whom you study, while assignment concerns which condition they receive.
An experiment still needs careful execution. Small groups may differ by chance, participants may not follow their assignments, outcomes may be missing, and the intervention may include several changes at once. Report the estimated effect and its uncertainty. Also distinguish evidence about the people and conditions studied from claims about every possible setting.
Do not make the coefficient answer the wrong question
Pearson's correlation coefficient, often written r, summarizes the direction and strength of a linear relationship between two quantitative variables. Its range is from −1 to 1. Penn State's regression materials emphasize that a value near zero does not rule out a curved relationship. A scatterplot can reveal features that one number hides.
A correlation of 0.9 also does not mean that changing one variable will produce a 90 percent change in the other, or that 90 percent of outcomes are caused by it. The coefficient has no measurement units and does not identify a causal direction. An impressive-looking number cannot substitute for a study design.
Make a useful conclusion without overclaiming
For the fictional class, a defensible statement is: tutoring attendees scored higher on average, but the difference disappeared within the two prior-preparation groups. That statement preserves the observation and its limitation. It avoids declaring either that tutoring caused the advantage or that tutoring can never help.
Associations remain useful for generating hypotheses, identifying groups that need attention, and building predictions that are separately validated. Causal inference can also use carefully designed observational evidence when experiments are infeasible, but that requires explicit assumptions and specialized methods. As a reader, ask which conclusion the study design actually supports, then keep the wording within that boundary.
Common questions
Does a very large sample prove causation?
No. More observations can improve precision, but they do not automatically remove confounding, selection bias, or measurement problems. A precisely estimated association may still have several causal explanations.
Is correlation useless without causation?
No. It can guide investigation or prediction. The key is not to assume that changing a predictor will necessarily change the outcome in the way an observed relationship suggests.
Can causation be studied without a randomized experiment?
Yes. Researchers use designs such as natural experiments and other causal inference methods. Their conclusions depend on the design and stated assumptions, so an observational label alone neither establishes nor automatically rules out a causal interpretation.
Sources & further reading
These references explain the underlying concepts. Examples on this page are illustrative; the source organizations do not endorse this site.
AI-assisted explanation. Read our editorial policy for scope and limitations. Found an error? Send a correction.