What Is a Null Hypothesis? Assuming Nothing Is Happening and Seeing If That Holds
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
Statistical testing works by supposing there is no effect and asking how surprising the data would be if that were true. The logic is indirect, it is widely misunderstood, and the misunderstandings account for a great deal of misreported research.
The structure of the test
A null hypothesis states that there is no difference, no association or no effect, and it is the proposition the test actually addresses. Data are collected and a statistic is computed, then the probability of obtaining a result at least as extreme as the one observed is calculated on the assumption that the null hypothesis is true. That probability is the p-value. If it falls below a chosen threshold the null is rejected, and if it does not the null is not rejected, which is deliberately not the same as accepting it. The logic is a proof by contradiction weakened for uncertainty, since a result very unlikely under the null counts against the null. Everything awkward about the procedure follows from its indirectness, since it never evaluates the hypothesis anyone is interested in.
What the p-value is not
The misinterpretations are specific, common and consequential:
- •It is not the probability that the null hypothesis is true, which is a different quantity requiring different assumptions
- •It is not the probability that the result occurred by chance
- •It is not the probability that the finding will replicate
- •One minus it is not the probability that the alternative hypothesis is true
- •It says nothing about the size of an effect, since a tiny effect in a large study gives a small value
- •A value above the threshold is not evidence of no effect, only absence of sufficient evidence against the null
Where the threshold came from
The conventional cut-off of one in twenty was a suggestion rather than a principle. Fisher proposed it as a convenient level at which a result might be worth a second look, explicitly as a rough guide, and he expected judgement rather than mechanical application. A separate framework developed by Neyman and Pearson treated testing as a decision procedure with specified error rates, which is a coherent but different scheme, and the two were subsequently merged into a hybrid that neither party endorsed and that is what is now taught in most introductory courses. The threshold's arbitrariness is visible in the fact that a result just inside it and one just outside are treated categorically differently despite being nearly identical, which has no statistical justification and considerable influence on what gets published.
The other kind of error
The framework defines two errors and attention concentrates on the wrong one. A false positive means rejecting the null when it is true, concluding there is an effect when there is none, and the threshold controls its rate directly. A false negative means failing to reject a false null, missing a real effect, and its rate depends on the size of the study and the size of the effect, which is what statistical power measures. The two trade against each other, since a stricter threshold reduces false positives and increases false negatives, and choosing between them is a judgement about which error costs more in the situation at hand. Most published discussion concerns false positives, while underpowered studies producing false negatives are arguably the larger problem in several fields, since they waste resources and leave real effects unestablished.
What is being done about it
The problems are recognised and the responses are varied. Reporting effect sizes and confidence intervals alongside or instead of test results is now standard guidance and gives readers the magnitude and the precision, which is what they actually need. Some journals have banned significance testing outright. Proposals to lower the threshold have been made and criticised for treating an arbitrary line as the problem rather than the reliance on any line. Bayesian methods evaluate hypotheses directly against each other and require stating prior beliefs, which is a genuine alternative with its own difficulties. Preregistration prevents the flexible analysis that produces significant results from noise. Estimation rather than testing is advocated as a general reorientation. The common thread is moving away from a binary verdict towards reporting how large an effect is and how well it was measured.
The takeaway
The test supposes no effect and computes how surprising the data would be under that supposition, which never evaluates the hypothesis of interest. The resulting probability is not the chance the null is true, not the chance the result came from chance and not the chance it will replicate. The threshold was a rough suggestion, and the current direction is reporting effect size and precision instead of a verdict.