What Is P-Hacking? Searching a Dataset Until Something Looks Significant
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
Analysing data in many ways and reporting only the analysis that produced a significant result generates findings that are not there. The practice is mostly not deliberate deception, which is what makes it difficult to eliminate.
How it happens
A dataset supports many defensible analyses, and each has some chance of producing a significant result by luck alone. A researcher who tries several and reports the one that worked has effectively run many tests while reporting one, which inflates the true error rate far above the nominal threshold. The crucial point is that each individual decision is usually defensible, since excluding an outlier, choosing one measure over another, adding a control variable or analysing a subgroup are all legitimate choices in isolation, and the problem arises from making them after seeing which way they push the result. That is why the practice frequently involves no awareness of wrongdoing, since the researcher can justify every step and is not lying about any of them, while the reported result is nonetheless unreliable.
The specific moves
Researchers surveyed anonymously admit to a recognisable list:
- •Collecting more data after checking whether the result is significant, and stopping when it is
- •Excluding participants or observations after seeing how the exclusion affects the outcome
- •Reporting only the outcome measures that worked, from several collected
- •Adding or removing control variables until the result reaches the threshold
- •Splitting the sample and reporting the subgroup where an effect appeared
- •Deciding after the fact which hypothesis was being tested, which makes an exploratory finding look confirmatory
How it is detected
Several methods identify the signature in a body of literature rather than in a single paper. The distribution of reported p-values across many studies can be examined, and a genuine effect produces a distribution with more very small values while a p-hacked literature shows a suspicious clustering just below the conventional threshold, which has been found repeatedly. Comparing preregistered studies with unregistered ones in the same field shows systematically smaller effects in the registered set. Checking whether reported statistics are internally consistent catches errors and some fabrication, and automated tools doing that have found inconsistencies in a substantial share of published papers. Examining whether the analysis reported matches the one preregistered catches undisclosed changes directly. None of these identifies an individual researcher's intent, and all of them establish that the practice is widespread.
The garden of forking paths
A subtler version of the problem operates without anyone trying multiple analyses at all. A researcher analysing a dataset once, having decided the approach after seeing the data, has still made choices that would have been different had the data looked otherwise, and the resulting flexibility inflates the error rate exactly as multiple testing would even though only one test was run. That is why the problem cannot be addressed by asking researchers to report every analysis they performed, since the analyses they would have performed on other data are the issue. Preregistration handles it by fixing the choices before the data exist, which is the only clean solution. The observation also explains why a researcher can honestly deny having tried many things while the result remains unreliable, and it is the version of the problem that is hardest to communicate.
What reduces it
The effective responses are structural rather than exhortative. Preregistration fixes the hypothesis, the sample size and the analysis before the data exist, which removes the flexibility entirely for the registered analysis while permitting exploratory work to be reported as such. Registered reports go further by having the study accepted on its design. Multiverse analysis reports the result under every defensible combination of choices, which shows how much the conclusion depends on the decisions rather than hiding that. Blinding the analyst to group labels until the analysis is fixed prevents the choices being steered. Larger samples reduce the noise that the practice exploits. And distinguishing exploratory from confirmatory work honestly, rather than presenting the first as the second, removes most of the problem at the cost of admitting that a finding is preliminary.
The takeaway
Many defensible analyses each have some chance of a significant result by luck, so trying several and reporting the one that worked inflates the error rate while every individual step remains justifiable. That is why it is mostly not deliberate. Clustering of reported values just below the threshold across a literature is the detectable signature, and preregistration removes the flexibility it depends on.