Inference and its limits
In this lesson
By the end, you’ll be able to
- State precisely what a p-value is and is not
- Separate statistical significance from effect size, and read a confidence interval as a compatibility range
- Recognize the garden of forking paths and why one result selected from many is weak evidence
Sign in to save your progress
You can keep reading without an account, but completed lessons won’t be saved.
Statistics, in a paper, is not mathematics for its own sake; it is a discipline of honesty under uncertainty. It gives you tools to say how much the data should move your belief, and, just as importantly, how easily you can fool yourself. Almost every tool in this lesson is more often misread than read. The goal is not to compute; it is to interpret without deceiving yourself or your reader.
A p-value answers one narrow question: if there were truly no effect, how surprising would data at least this extreme be? A small p-value means the data sit awkwardly with the null hypothesis; nothing more. It is not the probability that the null is true, not the probability your result is a fluke, and not one-minus-the-probability the effect is real. Those readings all smuggle in the very thing a p-value cannot give you: the probability of a hypothesis. And 'p < 0.05' is a convention, not a law of nature; it is a threshold someone chose, not a boundary between real and unreal.
Exercise
A study reports p = 0.03. Which statement is a correct reading of that number?
A p-value is computed under one assumption. Which one, and about whose truth?
Significance and size are different questions. A p-value tells you whether an effect is distinguishable from zero; an effect size tells you how big it is; a confidence interval tells you the range of effects reasonably compatible with your data. Sample size drives significance, so in a big enough study a trivial difference can reach p < 0.001 and still not matter to anyone. Always ask for the magnitude and its interval, and read the interval as a compatibility range, not as a 95% probability that the truth sits inside this particular one.
Exercise
A trial reports a mean improvement of 1.2 points, 95% CI [0.1, 2.3]. Which reading is most defensible?
Read the interval as a range of compatible effects, and notice how close its lower end is to zero.
Now the trap that undoes more findings than any other. Every analysis involves choices: which subgroup, which outcome, which cutoff, which covariates. Each defensible choice is a fork, and a researcher who decides after seeing the data can wander down whichever path reaches significance, often without any sense of wrongdoing. This is the garden of forking paths, and its arithmetic is brutal: run twenty independent tests on data where nothing is real and you have about a 64% chance of at least one 'significant' result. A finding cherry-picked from many comparisons is not strong evidence; it is the expected noise. The discipline that defeats it is to fix your one primary analysis before the data arrive, and to label everything else as exploration.
Signature interactive
Forking Paths Simulator
Here are 100 studies of pure noise, in every one, there is truly no effect at all. Choose how many analysis choices each team tries (a subgroup here, a different outcome there), and watch how many of them still find something “significant.”
22 of 100 studiesfound at least one “significant” result, in data where nothing is real.
Probability theory predicts about 23% (that is 1 − 0.955). With 5 tests, a false positive is already common.
Each orange square is a false alarm, a team that would, in good faith, write up a real-looking finding. This is why a result selected from many comparisons is weak evidence, and why deciding your one primary analysis before seeing the data is the fix.
Written response
Lab (constructed study). A team measured a new study app against 18 outcomes (grades in each of six subjects, three well-being scales, sleep, attendance, and more). One, chemistry grades, reached p = 0.04; the abstract's headline is 'the app significantly improves academic performance.' Write a short critique: what is wrong with the inference, and what two things would you need to see before believing the claim?
How many tests? What is the chance of at least one 'hit' if nothing is real? Then: effect size, interval, correction or replication.
0/70 words
Checkpoint · item 1 of 5
A study of 500,000 people finds that a supplement lowers a risk score by 0.02 points, p < 0.001. What is the right takeaway?
What does a very large sample do to significance for even a tiny difference?
Reflection