What Is the Replication Crisis? Findings That Do Not Hold Up Twice
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
Science relies on results being repeatable, and beginning around 2011 a series of large coordinated projects tried repeating well-known published findings and failed to reproduce a substantial fraction of them. The problem turned out not to be fraud in most cases but a set of ordinary practices that everyone was using and that reliably generate results which are not there.
What was found
Several large efforts produced consistent and uncomfortable numbers. A collaboration reproducing one hundred psychology studies published in leading journals obtained statistically significant results in well under half, and the effects that did replicate were on average about half the size originally reported. Comparable projects in cancer biology, economics and social science found similar patterns, with replication rates varying by field and rarely approaching what anyone expected. Industry laboratories reported privately and then publicly that they could not reproduce a majority of published preclinical findings they attempted to build drug programmes on. Surveys of researchers themselves found large proportions reporting they had failed to reproduce someone else's work and, notably, their own. Several specific famous results, including some widely taught, failed to hold up under high-powered replication, and the field-wide response was initially defensive before becoming, in most disciplines, a serious reform effort.
How ordinary practice produces false results
The causes are mostly structural rather than dishonest, which is what makes them hard to fix:
- •Low statistical power, since small samples detect only large effects, and a small study that does find an effect has almost certainly overestimated its size
- •Researcher degrees of freedom, meaning the many defensible choices about excluding outliers, which measures to use and which covariates to include, each of which can be made after seeing the data
- •Analysing until something works, sometimes called p-hacking, which need not be deliberate and which reliably produces a publishable result from noise
- •Hypothesising after results are known, presenting an exploratory finding as if it had been predicted, which removes the protection that prediction provides
- •Publication bias, since journals preferred novel positive findings and rarely published null results or replications, so the literature shows the successes and hides the failures
- •Incentives, since careers depend on publication counts and prestigious placements rather than on whether findings hold
- •The statistical threshold itself, since a fixed significance cut-off treats a continuous measure of evidence as a binary verdict and encourages exactly the behaviour above
What has changed
The reforms adopted are specific and many are now standard in the affected fields. Preregistration requires researchers to publish their hypothesis, sample size and analysis plan before collecting data, which removes the flexibility that produces false positives, and registered reports go further by having journals accept a study on the basis of its design before results exist, which eliminates publication bias by construction. Power analysis before data collection and substantially larger samples have become expected. Data and code sharing allows others to check analyses directly. Multi-laboratory collaborations run the same protocol across many sites, producing estimates far more reliable than any single study. Journals have created formats for null results and for replications, which were previously close to unpublishable. Some fields have lowered the significance threshold or moved towards reporting effect sizes and intervals rather than binary verdicts. Early evidence suggests preregistered studies report positive results far less often, which is what the reform was supposed to do.
What it does and does not mean
Two misreadings are common and both are worth resisting. The first treats the crisis as evidence that science cannot be trusted, which does not follow, since the crisis was discovered and quantified by scientists using scientific methods and has produced substantial correction, which is the process working rather than failing. The second treats it as confined to psychology, when replication problems have been documented across biomedicine, economics, ecology and beyond, and the underlying incentives are shared. The more useful conclusions are practical: a single study, particularly a small and surprising one, is weak evidence and should be read as such; effect sizes in early reports are systematically inflated; and the fields with the strongest reforms are now producing more trustworthy findings than those without. It also implies something about reading science journalism, since the results most likely to be reported are the novel and surprising ones, which are precisely the ones least likely to replicate.
The takeaway
Large coordinated projects from 2011 onward reproduced well under half of published findings in several fields, with surviving effects around half the original size. The causes were ordinary practice rather than fraud: small samples, flexible analysis choices made after seeing data, presenting exploratory results as predictions, and journals publishing only positive novel findings. Reforms include preregistration, registered reports, data sharing and multi-site collaborations, and preregistered studies report far fewer positive results.