"A Study Showed That…": How to Actually Read a Paper
Four words that end most arguments
"A study showed that…" — and the conversation stops, because who argues with a study? Studies are science, science is facts, facts win.
Except "a study showed that" is one of the least informative sentences in the language. There are millions of studies. They range from careful, replicated work to a survey of nineteen undergraduates that will never be reproduced. Some are excellent. Many are wrong — not fraudulent, just underpowered, badly designed, or unlucky. And by the time a finding reaches you, it's usually passed through a press office and a headline writer, each of whom made it stronger and simpler than the paper ever claimed.
So the skill isn't trusting studies or distrusting them. It's being able to open one and judge it — which most people, including most people who cite studies, cannot do. This is an L3 lesson because it's genuinely hard, and it's one of the highest-leverage things in the whole Data pillar: once you can read a paper, "a study showed that" stops being a trump card and becomes a claim you can check.
Why this is hard, and why it's worth it
Papers are written to be published and to satisfy reviewers, not to be understood by you. They're dense, hedged, full of jargon, and structured in a way that buries the two things you most want (does it hold up, and how much does it actually show) inside things you care about less.
But there's a secret that makes it tractable: you do not read a paper front to back, and you do not need to understand every method. There's an order that gets you to a judgement fast, and a short list of questions that catch most of the ways a paper misleads. That's what this lesson is.
The stakes are real. Ioannidis's famous argument — that a large fraction of published findings are false — and the reproducibility work that followed, where a majority of some fields' headline results failed to replicate, aren't reasons to dismiss science. They're reasons to read individual claims critically, because "it was published" turns out to guarantee much less than people assume.
Mechanism 1 — Read it in the right order
Front-to-back is the slow, confusing way. Here's the order that gets you to a verdict:
1. TITLE + ABSTRACT what do they claim to have found?
2. Skip to METHODS what did they ACTUALLY do? (the real paper)
3. FIGURES + TABLES the results, before their spin on them
4. LIMITATIONS what do THEY admit is weak? (usually near the end)
5. Now the discussion their interpretation — read last, skeptically
6. Funding + conflicts who paid, and who benefits from this result?The order matters because the abstract and discussion are where the paper argues for itself — they put the finding in its best light. The methods and figures are where the paper is what it is. Reading the argument before the evidence lets the argument frame the evidence. Read the evidence first.
Most people read the abstract and stop. The abstract is a trailer, edited to sell the film.
Mechanism 2 — The six questions
Run these on any study. They catch most of the ways a real paper overclaims.
1. How many people/things? (sample size)
The single fastest filter. A finding from 20 people is a hint; a finding from 20,000 is closer to evidence. Small samples produce dramatic, unstable results that vanish on repetition — the smaller the study, the more its result is driven by noise and luck.
Red flag: a big, confident claim from a tiny sample. "Chocolate improves memory (n=15)" is not a reason to eat chocolate.
2. Was there a comparison group? (control)
To know if something works, you compare people who got it against people who didn't. No control group means you can't separate the effect from everything else that was happening. "People who did our programme improved" is meaningless without "…more than people who didn't."
The gold standard is a randomised controlled trial — people assigned to groups by chance, so the groups don't differ systematically. Randomisation is what lets you attribute the difference to the treatment rather than to who chose it.
Red flag: no control group, or groups that people sorted themselves into.
3. Correlation or causation? (what kind of study)
The distinction that trips up most headlines (DT-03 is the whole lesson). An observational study watches what happens — it can show two things move together, never that one causes the other. An experimental study intervenes and can, carefully, support causation.
Watch the verbs. Papers usually hedge honestly — "is associated with", "is linked to". Then the press release turns "associated with" into "causes", and the headline drops the hedge entirely.
Red flag: an observational study being described (by anyone) with causal language.
4. Is the effect big, or just "significant"? (effect size vs p-value)
This one is subtle and important. "Statistically significant" does not mean "large" or "important." It means the result probably isn't pure chance — that's all. A tobacco-sized study can find a statistically significant effect so tiny it means nothing in real life.
Two numbers to separate:
- The p-value answers "could this be a fluke?" (roughly). Convention is p < 0.05, which is a low bar and widely misused.
- The effect size answers "how much?" — the actual magnitude. This is the one that matters for whether you should care.
Red flag: a paper (or article) trumpeting significance while never telling you how big the effect was. "Significant" without a magnitude is hiding something. (Full treatment: DT-06.)
5. Who paid, and who benefits?
Funding doesn't automatically invalidate a study, but it bends the odds. Research funded by a party with a stake in the outcome is, on average, more likely to find what that party wanted — through a hundred small, often unconscious choices, not usually through fraud.
Where to look: the "funding", "conflicts of interest", or "acknowledgements" sections. A supplement study funded by the supplement maker isn't worthless, but it earns extra scrutiny.
Red flag: the funder benefits directly from the result, especially combined with any of the above.
6. Has anyone repeated it? (replication)
A single study is a data point, not a fact. The most important question of all and the one most often skipped. A finding becomes trustworthy when other teams, independently, get the same result. A lone, surprising, never-repeated study — however exciting — is exactly the kind that tends to evaporate.
Red flag: a big claim resting on one study, especially a surprising one, especially recent, before anyone has tried to reproduce it. The more a finding contradicts everything known before, the more replication it needs before you believe it — extraordinary claims, ordinary evidence, is how surprises turn out wrong.
Mechanism 3 — The distortion chain
Even a good paper reaches you damaged, because it travels through a pipeline that strengthens it at every step:
The paper "X was associated with a modest increase in Y
in this sample; further study is needed."
↓ press office
Press release "Scientists discover X boosts Y!"
↓ journalist
Headline "X is the secret to Y"
↓ social media
What you see "X CAUSES Y 🤯 (a study showed)"Each step drops a hedge, adds certainty, and swaps association for causation. By the time it's a headline or a post, it can be the near-opposite of what the researchers wrote.
The practical consequence: go upstream. When a claim matters, find the actual paper (its title is usually in the article, or search the finding + "study"). Read its abstract, not the headline's version. The gap between the two is frequently the whole story — and closing it is most of what this lesson is for.
What this means for you
- "A study showed that" is the beginning of a question, not the end of an argument. Which study? How big? Repeated?
- One study is a data point. Trust builds from replication, not from publication.
- "Significant" ≠ "big" ≠ "important." Always ask for the effect size, not just whether an effect exists.
- Read methods before discussion. The methods are what happened; the discussion is the sales pitch.
- Go upstream to the paper when it matters. The headline is the finding after several people made it stronger than it was.
- This is not "don't trust science." It's the opposite: science works because claims get checked, and now you can be one of the people checking rather than one of the people forwarding.
Try it: evaluate one real paper (60 min)
- Find a headline making a scientific claim you find interesting — health, psychology, diet, technology, anything.
- Find the actual paper it's based on. (Look for the study name in the article; search the claim plus "study" or use a scholar search. Many are open-access; an abstract is enough to start.)
- Read it in the order from Mechanism 1 — methods and figures before discussion.
- Answer the six questions in writing:
1. SAMPLE SIZE: n = ...... (hint / evidence?)
2. CONTROL GROUP? yes / no / self-selected
3. CORRELATION or CAUSATION? observational / experimental —
and which language did the paper use vs the headline?
4. EFFECT SIZE: how big, in real terms? .............
(did the article tell you, or only "significant"?)
5. FUNDING / CONFLICTS: who paid? who benefits? ............
6. REPLICATED? one study, or repeated? ............- Write a verdict: does the paper support the headline? In one paragraph, say what the study actually shows, and where the headline went beyond it.
✅ Finish check: one real paper run through all six questions in writing, plus a one-paragraph verdict comparing what the study shows to what the headline claimed.
Summary card
- "A study showed that" is nearly uninformative. Millions of studies exist, of every quality. Ask which one, how good.
- Read in order: title/abstract → methods → figures → limitations → discussion (last) → funding. The methods are the real paper; the discussion is the pitch.
- The six questions: sample size · control group · correlation vs causation · effect size (not just significance) · funding/conflicts · replication.
- "Statistically significant" ≠ large ≠ important. Demand the effect size.
- One study is a data point. Trust comes from independent replication.
- Watch the verbs: "associated with" is not "causes." Headlines routinely swap them.
- The distortion chain strengthens a claim at every step. Go upstream to the paper.
- Extraordinary claims need extraordinary replication before you believe them.
- This is how science is supposed to be read — critically, not credulously.
Sources
- Goldacre, B. — Bad Science, 2008; Bad Pharma, 2012
- Ioannidis, J. — Why Most Published Research Findings Are False, 2005
- Greenhalgh, T. — How to Read a Paper, 2019
- Open Science Collaboration — Estimating the Reproducibility of Psychological Science, 2015
Next lesson: DT-06 — p-values, Significance and Effect Size (L3) Related: DT-01 Reading Data · DT-02 How Charts Lie · DT-03 Correlation Is Not Causation · DT-09 Verifying a Source · AI-07 Hallucination Hunting Path: Hard to Fool — 6/6