Inference and its limits

Lesson 8 40 minEvidence

In this lesson

By the end, you’ll be able to

  • State precisely what a p-value is and is not
  • Separate statistical significance from effect size, and read a confidence interval as a compatibility range
  • Recognize the garden of forking paths and why one result selected from many is weak evidence

Sign in to save your progress

You can keep reading without an account, but completed lessons won’t be saved.

Sign in

Statistics, in a paper, is not mathematics for its own sake; it is a discipline of honesty under uncertainty. It gives you tools to say how much the data should move your belief, and, just as importantly, how easily you can fool yourself. Almost every tool in this lesson is more often misread than read. The goal is not to compute; it is to interpret without deceiving yourself or your reader.

A p-value answers one narrow question: if there were truly no effect, how surprising would data at least this extreme be? A small p-value means the data sit awkwardly with the null hypothesis; nothing more. It is not the probability that the null is true, not the probability your result is a fluke, and not one-minus-the-probability the effect is real. Those readings all smuggle in the very thing a p-value cannot give you: the probability of a hypothesis. And 'p < 0.05' is a convention, not a law of nature; it is a threshold someone chose, not a boundary between real and unreal.

Exercise

A study reports p = 0.03. Which statement is a correct reading of that number?

A p-value is computed under one assumption. Which one, and about whose truth?

Answer options
How sure?

Significance and size are different questions. A p-value tells you whether an effect is distinguishable from zero; an effect size tells you how big it is; a confidence interval tells you the range of effects reasonably compatible with your data. Sample size drives significance, so in a big enough study a trivial difference can reach p < 0.001 and still not matter to anyone. Always ask for the magnitude and its interval, and read the interval as a compatibility range, not as a 95% probability that the truth sits inside this particular one.

Exercise

A trial reports a mean improvement of 1.2 points, 95% CI [0.1, 2.3]. Which reading is most defensible?

Read the interval as a range of compatible effects, and notice how close its lower end is to zero.

Answer options
How sure?

Now the trap that undoes more findings than any other. Every analysis involves choices: which subgroup, which outcome, which cutoff, which covariates. Each defensible choice is a fork, and a researcher who decides after seeing the data can wander down whichever path reaches significance, often without any sense of wrongdoing. This is the garden of forking paths, and its arithmetic is brutal: run twenty independent tests on data where nothing is real and you have about a 64% chance of at least one 'significant' result. A finding cherry-picked from many comparisons is not strong evidence; it is the expected noise. The discipline that defeats it is to fix your one primary analysis before the data arrive, and to label everything else as exploration.

one datasetchoice 1choice 2choice 3subgroup? outcome? cutoff?p = n.s.p = n.s.p = n.s.p = 0.03 ✓ reportedp = n.s.p = n.s.p = n.s.
One dataset, many defensible analysis choices, each a branch ending in its own test. Report only the branch that reached p < 0.05 and you have manufactured a finding from noise, without ever telling a conscious lie.

Signature interactive

Forking Paths Simulator

Here are 100 studies of pure noise, in every one, there is truly no effect at all. Choose how many analysis choices each team tries (a subgroup here, a different outcome there), and watch how many of them still find something “significant.”

Analysis choices per study

22 of 100 studiesfound at least one “significant” result, in data where nothing is real.

Probability theory predicts about 23% (that is 1 − 0.955). With 5 tests, a false positive is already common.

Each orange square is a false alarm, a team that would, in good faith, write up a real-looking finding. This is why a result selected from many comparisons is weak evidence, and why deciding your one primary analysis before seeing the data is the fix.

Written response

Lab (constructed study). A team measured a new study app against 18 outcomes (grades in each of six subjects, three well-being scales, sleep, attendance, and more). One, chemistry grades, reached p = 0.04; the abstract's headline is 'the app significantly improves academic performance.' Write a short critique: what is wrong with the inference, and what two things would you need to see before believing the claim?

How many tests? What is the chance of at least one 'hit' if nothing is real? Then: effect size, interval, correction or replication.

0/70 words

How sure?

Checkpoint · item 1 of 5

A study of 500,000 people finds that a supplement lowers a risk score by 0.02 points, p < 0.001. What is the right takeaway?

What does a very large sample do to significance for even a tiny difference?

How sure?

Reflection