Correlation vs Causation — Explained with Real Studies

Correlation vs Causation

Learn With Examples

Two things moving together is a fact. One thing causing the other is a claim. The gap between them has produced bad headlines, wasted billions in research money, and in a few famous cases, medical advice that actively harmed people.

Reading time13 min
LevelBeginner friendly
Includes5 real studies

In the 1990s, if you were a woman approaching menopause, your doctor may well have suggested hormone replacement therapy — not only for menopausal symptoms, but to protect your heart. The evidence looked strong. Large, careful observational studies following tens of thousands of nurses found that women taking HRT had substantially less coronary heart disease than women who didn’t.

Then, in 2002, a large randomised trial tested it properly. The heart protection wasn’t there. In the group studied, the therapy was associated with increased risks rather than reduced ones, and the trial arm was stopped early.

What went wrong? Nothing, statistically. The correlation was real. Women on HRT genuinely did have fewer heart attacks. But they were also, on average, wealthier, better educated, more likely to exercise, and more likely to see a doctor regularly. The therapy wasn’t protecting their hearts. Their circumstances were — and those same circumstances made them more likely to be prescribed HRT in the first place.

That’s the entire subject of this article, in one real example. Correlation is a measurement. Causation is an explanation. And the distance between them is where most misleading science reporting lives.

The core distinction

Two sentences that are not the same

Correlation: when one thing changes, another tends to change too. It’s an observation about patterns in data. It can be measured, quantified, and it’s often perfectly real.

Causation: changing the first thing makes the second thing change. It’s a claim about mechanism — and it requires far more evidence than a pattern.

Every correlation has at least four possible explanations, and only one of them is “A causes B”. The skill worth building is naming the other three.

First, what correlation actually measures

Correlation is usually expressed as a number between −1 and +1, written r. It tells you two things: the direction of the relationship and how tightly the points cluster around it. It tells you nothing whatsoever about why.

r = 0.0No relationship
r = 0.45Weak / noisy
r = 0.92Strong positive
r = -0.85Strong negative

Two properties matter here. First, a strong correlation is not stronger evidence of causation than a weak one — it’s just a tighter pattern. An r of 0.95 between two unrelated things is still a coincidence, merely a tidier-looking one. Second, correlation as usually measured only detects straight-line relationships. A real, powerful, curved relationship can produce an r near zero.

The number that fooled everyone. In 2012, a paper in a major medical journal noted that a country’s chocolate consumption correlated with its number of Nobel laureates per capita at roughly r = 0.79 — a strikingly strong figure. The author intended it partly as a joke about exactly this problem. It was reported widely and sincerely anyway. National wealth explains both: rich countries buy more chocolate and fund more research.

The five explanations for any correlation

When you see “X is linked to Y”, these are the candidates. Tap through them — each has a real example attached, and getting fluent in the list is most of the skill.

Why might two things move together?

tap an explanation
Pure chance

Coincidence

Compare enough variables and some will line up beautifully by luck alone. This isn’t a rare accident — it’s a mathematical certainty. Test a thousand random pairs at the usual significance threshold and roughly fifty will look “significant” with nothing behind them.

Real example: a 2000 paper in a teaching journal cheerfully demonstrated a statistically significant correlation between stork populations and human birth rates across European countries. The point was satire, and the mechanism is mundane: larger countries have more storks and more babies.

How to spot it: ask whether anyone predicted this relationship before looking. A pattern found by trawling data needs replication in fresh data before it means anything.

Most common culprit

A third variable (confounding)

Something else drives both. This is the explanation behind the majority of misleading health and social science headlines, and it’s the reason the HRT story went wrong.

Real example: a 1999 study reported that infants who slept with a night light were far more likely to become short-sighted. It made headlines worldwide. Follow-up studies the next year found the association vanished once parental short-sightedness was accounted for. Short-sighted parents were both more likely to use night lights and more likely to pass on the genes. The light was innocent.

How to spot it: ask who ends up in each group and why. If the groups differ in ways other than the thing being studied, the third variable is probably doing the work.

Backwards

Reverse causation

The arrow points the other way. B causes A, not A causes B — and in cross-sectional data, both look identical.

Real example: studies repeatedly find that people who skip breakfast weigh more, which produced years of “eat breakfast to lose weight” advice. When researchers actually randomised people to eat or skip breakfast, weight outcomes barely differed. A plausible reading is partly reverse: people already trying to lose weight skip meals.

Another: patients treated in hospital die at higher rates than people who stay home. Hospitals are not the danger — being seriously ill sends you there.

How to spot it: ask which came first. If the data can’t tell you, neither can the headline.

Who got counted

Selection bias

The relationship is an artefact of who ended up in the sample. Nothing is wrong with the analysis; the group being analysed was assembled in a skewed way.

Real example: the long-held belief that moderate drinking protects the heart rests heavily on comparisons with non-drinkers. But the “never drinks” group quietly includes people who stopped because they were already unwell — the “sick quitter” effect. Studies that separate lifelong abstainers from former drinkers, and genetic analyses that sidestep lifestyle entirely, find the protective effect largely disappears.

How to spot it: ask who is missing from the data, and why they’re missing.

Sometimes true

It really is causal

Sometimes A does cause B, and the correlation is the first clue. Dismissing every correlation is as lazy as accepting every one — the job is to work out which is which.

Real example: smoking and lung cancer. When the link emerged in the 1950s, sceptics argued exactly what this article argues: correlation isn’t causation, perhaps some genetic factor caused both the smoking habit and the cancer. That objection was answered not by one study but by a mountain of converging evidence — which is what real causal claims look like.

How to spot it: look for a dose-response pattern, a plausible mechanism, consistency across very different populations, and correct time ordering.

The question is never “is this correlation real?” It usually is. The question is “what else would produce this exact pattern?”

Five real studies, and what actually happened

These aren’t hypotheticals. Each is a case where a genuine correlation led somewhere — sometimes to a correction, once to a Nobel-worthy public health victory.

Case studies

tap a case
Confounding

Hormone therapy and heart disease

What observation suggestedLarge cohort studies through the 1980s and 90s found women on hormone replacement therapy had markedly less coronary heart disease. Guidelines and prescribing practice followed.
What the trial foundThe Women’s Health Initiative randomised trial, reported in 2002, found no cardiac protection. Risks in the studied population went the wrong way and the arm was halted early.

The lesson: the women taking HRT were healthier before they took anything — a pattern researchers now call the healthy-user effect. Careful statistical adjustment did not rescue the finding, because you can only adjust for confounders you thought to measure.

Confounding

Beta-carotene and lung cancer

What observation suggestedPeople eating diets rich in beta-carotene — carrots, leafy greens — had noticeably lower rates of lung cancer. Supplements looked like an obvious public health win.
What the trials foundTwo large randomised trials in the 1990s tested beta-carotene supplements in high-risk smokers. Lung cancer rates went up in the supplement groups, and one trial was stopped early.

The lesson: eating vegetables and swallowing an isolated compound are not the same intervention. The original correlation was probably tracking a whole dietary and lifestyle pattern, not the nutrient. This is one of the clearest cases where acting on a correlation caused measurable harm.

Confounding

Night lights and short-sightedness

What the 1999 study foundChildren who had slept with a night light or room light as infants were several times more likely to be short-sighted later. It was published in a top journal and reported around the world.
What followed in 2000Independent studies failed to replicate the effect once parental short-sightedness was included. Myopic parents were more likely to light the nursery — and more likely to pass on myopia.

The lesson: the fastest way to test a surprising finding is to ask what kind of household does the thing. Habits cluster with genetics, income and education, and any of those can be the real driver.

Selection bias

Moderate drinking and heart health

What decades of data suggestedA J-shaped curve: moderate drinkers appeared to have lower cardiovascular mortality than both heavy drinkers and non-drinkers. It became received wisdom, and a marketing gift.
What better methods suggestSeparating lifelong abstainers from people who quit for health reasons shrinks the benefit substantially. Genetic approaches, which avoid lifestyle confounding by design, find little support for a protective effect.

The lesson: the comparison group matters as much as the exposure group. If your baseline is quietly full of sick people, everything else looks healthy by contrast.

Genuinely causal

Smoking and lung cancer

The correlationStudies from the early 1950s, including a landmark study following British doctors for decades, found heavy smokers had dramatically elevated lung cancer rates.
Why it held upEnormous effect size, a clear dose-response gradient, risk falling after quitting, consistency across countries and study designs, animal and laboratory evidence, and an identified biological mechanism.

The lesson: this is the template for establishing causation without a randomised trial — you cannot ethically assign people to smoke. The criteria used to argue it, published in 1965, are still the standard checklist. Correlation was the starting point, not the conclusion.

Four corrections and one confirmation. That ratio is roughly what the history of nutritional and lifestyle epidemiology looks like — which is a good reason to treat single observational findings as questions rather than answers.

So how is causation established?

Not by one study. By methods that progressively strip away the alternative explanations. Here they are, weakest to strongest.

Anecdote
A story. Zero control over anything.
Cross-sectional
A snapshot. Can’t even establish which came first.
Cohort study
Follows people over time. Fixes time order, not confounding.
Natural experiment
Something outside the researcher’s control split the groups — a policy change, a lottery.
Randomised trial
Assignment by chance, so groups differ only by luck.
Replicated body of evidence
Many designs, many populations, one consistent answer.

Randomisation is the crucial step, and it’s worth understanding why it works. If you flip a coin to decide who gets the treatment, then income, diet, genetics, motivation and every other variable — including ones nobody has thought of — end up distributed roughly evenly between the groups. Statistical adjustment can only handle confounders you measured. Randomisation handles the ones you didn’t.

When you can’t randomise

You cannot assign people to smoke for thirty years, or to grow up poor. Researchers use other tools:

Natural experiments

A policy changes in one state and not its neighbour. Compare the two and the difference approximates a trial nobody ran.

Genetic approaches

Gene variants are allocated at conception, effectively at random. Comparing people by variant sidesteps lifestyle confounding — the method that undercut the drinking story.

Dose-response

If more exposure reliably means more effect, coincidence gets much harder to argue.

Triangulation

Several methods with different weaknesses pointing the same way. Each design’s blind spot is covered by another’s strength.

Reading a headline in thirty seconds

Six questions. You don’t need statistical training for any of them.

Ask thisWhy it mattersWarning sign
Trial or observation?Randomisation is what removes unknown confounders“Linked to”, “associated with”
Who is being compared?Groups may differ in a dozen unmeasured waysSelf-selected groups, volunteers
Could it run backwards?Reverse causation looks identical in a snapshotBoth measured at the same moment
Humans, and how many?Small or animal studies rarely justify adviceUnder a few hundred; mice
Relative or absolute risk?“Doubles your risk” can mean 1 in 100,000 to 2Percentages with no baseline
Has it replicated?A single striking finding is a hypothesis“New study finds”, “first evidence”

A quick vocabulary tell. Careful researchers write “associated with”, “linked to”, or “correlated with” precisely because they mean correlation. Headline writers then translate that into “causes”, “boosts”, “triggers” or “cures”. When the study says one thing and the headline says the other, trust the study — and notice that the gap is usually introduced after the science is finished.

Where this bites outside science

Business decisions

Customers using your new feature retain better — or your most committed customers were always the ones who tried new features.

Marketing spend

Ad spend correlates with sales. It also rises during the seasons when sales rise anyway. A holdout region tells you far more than a chart.

Education

Children in smaller classes score higher — and smaller classes cluster in wealthier districts with more of everything else.

Personal health

You started a supplement and felt better. So did the season, your sleep, and the illness that was ending anyway.

Check yourself

Five questions. Open each to check — the correct option is marked.

1. A study finds people who drink more coffee live longer. What’s the weakest conclusion to draw?
  • Coffee drinkers differ from non-drinkers in other ways
  • Coffee extends lifespan
  • The finding deserves a randomised test
  • Health may affect coffee habits, not the reverse

Observational data supports the association. The causal claim needs the alternatives ruled out first — and sick people often cut down on coffee, which is reverse causation.

2. In the night-light and myopia case, what was the third variable?
  • Screen time
  • Room size
  • Parental short-sightedness
  • The child’s age

Short-sighted parents were both more likely to use night lights and more likely to pass on myopia genetically. Account for the parents and the effect disappears.

3. Why does randomisation remove confounding so effectively?
  • It increases the sample size
  • It balances variables the researchers never even measured
  • It makes the correlation stronger
  • It removes measurement error

Statistical adjustment can only fix confounders you thought to record. Coin-flip assignment spreads the unknown ones evenly too.

4. “Hospital patients die more often than people at home.” What’s the flaw?
  • Coincidence
  • Reverse causation — illness sends people to hospital
  • The sample is too small
  • There’s no flaw

The arrow runs backwards. Being seriously ill causes hospital admission; the hospital didn’t cause the illness.

5. Why was smoking accepted as causing lung cancer without a randomised trial?
  • The correlation was very strong
  • A single large study settled it
  • Many independent lines of evidence converged — dose-response, mechanism, consistency, risk falling after quitting
  • Governments decided it

Strength alone was never sufficient — sceptics raised exactly the confounding objection. It was the convergence of many designs that closed the case.

Frequently asked questions

What is the difference between correlation and causation in simple terms?

Correlation means two things move together in the data. Causation means changing one actually produces a change in the other. Correlation is measured; causation is inferred, and it needs the alternative explanations ruled out first.

Can correlation ever prove causation?

Not on its own. But correlation plus dose-response, correct time ordering, a plausible mechanism, consistency across different populations and study designs, and preferably experimental evidence together can establish causation convincingly — as happened with smoking.

What is a confounding variable?

A third factor that influences both things being studied, creating a correlation between them that isn’t causal. Wealth confounds chocolate consumption and Nobel prizes; parental short-sightedness confounded night lights and myopia.

Why are randomised controlled trials considered the gold standard?

Because assignment by chance makes the groups comparable on everything, including factors nobody thought to measure. That’s the one thing statistical adjustment of observational data can never fully achieve.

Does “correlation is not causation” mean I should ignore observational studies?

No. Observational research is how most hypotheses start, and for many questions it’s the only ethical option. Treat a single observational finding as a good question rather than a settled answer, and give much more weight to results that replicate across different methods.

The takeaway

Correlation is a fact about data. Causation is a claim about the world. Moving from one to the other requires ruling out coincidence, third variables, reverse causation and selection bias — and the historical record shows how often that fails, even in careful hands, even in top journals.

The HRT story is worth keeping in mind precisely because nobody was careless. Good researchers, huge samples, real statistical rigour, and the answer was still wrong — because the women who took the therapy were different from the women who didn’t in ways the data never captured.

You don’t need to become a statistician. You need one reflex: when you read that X is linked to Y, pause and ask what else could produce that exact pattern. Usually you’ll think of something within about ten seconds. That reflex is worth more than any formula in this article.

correlation vs causationstatisticsconfoundingresearch methodscritical thinkingdata literacy

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *