What Is a P-Value? The Most Misunderstood Number in Science
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
It is not the probability that the hypothesis is true. It is not the probability that the result occurred by chance. It is not the probability that the finding will replicate. A p-value is the probability of observing data at least as extreme as what was observed, assuming the null hypothesis is true, and every popular restatement of it that drops that final clause is wrong in a way that has consequences.
What it actually says
The logic runs backwards from what most people assume. You begin by assuming there is no effect, which is the null hypothesis. You then ask: if that assumption were true, how surprising would this data be? The p-value quantifies that surprise, so a value of 0.03 means that data this extreme or more so would occur three percent of the time in a world where the effect does not exist. It says nothing directly about the world where the effect does exist, and it cannot, because it never considered that world. Converting it into a statement about the probability of the hypothesis requires knowing how likely the hypothesis was before the data, which is Bayes's theorem, and the p-value contains no such information. That is why a small p-value from a study of an implausible hypothesis is much weaker evidence than the same p-value from a plausible one, and why a field testing many unlikely ideas generates a great many false positives at any threshold.
Where 0.05 came from
The conventional threshold has no theoretical basis. Ronald Fisher, who introduced the measure in the 1920s, suggested 0.05 as a convenient round figure, wrote that it was a matter of personal judgement, and explicitly said that a single significant result should not be treated as settling anything, proposing instead that an experimenter should be able to produce the result reliably. The threshold hardened into a rule for reasons of convenience: printed tables of critical values were expensive to produce, so a few standard levels were tabulated and used, and journals then adopted them as a publication criterion. The consequence is a binary treatment of a continuous quantity, in which 0.049 and 0.051 are reported as categorically different when they are almost identical, and in which a result just above the line is frequently described as a trend toward significance, a phrase that means nothing.
The misinterpretations
A 2016 statement by the American Statistical Association, unusual because the organisation had never issued guidance on a specific statistical practice before, listed the errors explicitly:
- •A p-value does not measure the probability that the hypothesis under study is true
- •It does not measure the probability that the data were produced by chance alone
- •A threshold of 0.05 does not make a decision correct, and scientific conclusions should not be based on whether a value crosses a line
- •A p-value does not measure the size of an effect or its importance, since a trivially small effect measured in a very large sample produces a very small p-value
- •By itself it does not provide a good measure of evidence, since the same value means different things depending on the plausibility of the hypothesis and the design of the study
- •Non-significant does not mean no effect, since an underpowered study routinely fails to detect real effects, and absence of evidence is not evidence of absence
How the threshold gets gamed
Because publication depends on crossing the line, a set of practices has developed that inflate false positives without anyone lying. Collecting data until the result becomes significant and then stopping, testing many outcomes and reporting the ones that worked, dropping inconvenient participants, trying several analytical approaches and reporting the best, and forming the hypothesis after seeing the results while presenting it as prior are collectively known as p-hacking. Simmons, Nelson and Simonsohn demonstrated in 2011 that a modest combination of these practices raises the false positive rate from five percent to over sixty, and illustrated it by using real data and legitimate-looking analysis to prove that listening to a particular song made people younger. The garden of forking paths, described by Andrew Gelman, is the subtler version: a researcher making reasonable analytical choices in response to their data, with no intent to deceive, still effectively runs many analyses and reports one.
What is being done
The replication crisis, in which large coordinated projects found that a substantial proportion of published findings in psychology, cancer biology and economics did not reproduce, brought the issue to general attention and produced a set of reforms. Pre-registration commits to hypotheses and analysis before data collection, and registered reports go further by having journals accept a study on the strength of its design before results exist, which removes publication bias at the source. Reporting effect sizes and confidence intervals rather than only a significance verdict conveys magnitude and precision. Some journals have banned p-values outright. A widely discussed proposal to lower the threshold to 0.005 for claims of new discoveries drew a counterargument that the problem is the threshold mentality rather than its location. Bayesian methods, which do estimate the probability of hypotheses given the data and require stating prior beliefs, have gained ground. The consistent recommendation across all of it is that a p-value is one piece of information about a result, and that treating it as a verdict is the error.
The takeaway
A p-value is the probability of data at least this extreme assuming no effect exists, which says nothing directly about whether an effect exists and cannot without knowing how plausible it was beforehand. The 0.05 threshold was a convenience Fisher proposed and explicitly warned against treating as decisive. It measures neither the size nor the importance of an effect, non-significance does not mean nothing is there, and practices that select analyses after seeing data can raise the false positive rate from five percent to over sixty without any dishonesty.