← All articles
sciencereplicationresearchstatisticsSeptember 17, 20264 min read

What Is the Replication Crisis? Findings That Do Not Hold Up Twice

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

Science relies on results being repeatable, and beginning around 2011 a series of large coordinated projects tried repeating well-known published findings and failed to reproduce a substantial fraction of them. The problem turned out not to be fraud in most cases but a set of ordinary practices that everyone was using and that reliably generate results which are not there.

What was found

Several large efforts produced consistent and uncomfortable numbers. A collaboration reproducing one hundred psychology studies published in leading journals obtained statistically significant results in well under half, and the effects that did replicate were on average about half the size originally reported. Comparable projects in cancer biology, economics and social science found similar patterns, with replication rates varying by field and rarely approaching what anyone expected. Industry laboratories reported privately and then publicly that they could not reproduce a majority of published preclinical findings they attempted to build drug programmes on. Surveys of researchers themselves found large proportions reporting they had failed to reproduce someone else's work and, notably, their own. Several specific famous results, including some widely taught, failed to hold up under high-powered replication, and the field-wide response was initially defensive before becoming, in most disciplines, a serious reform effort.

How ordinary practice produces false results

The causes are mostly structural rather than dishonest, which is what makes them hard to fix:

  • Low statistical power, since small samples detect only large effects, and a small study that does find an effect has almost certainly overestimated its size
  • Researcher degrees of freedom, meaning the many defensible choices about excluding outliers, which measures to use and which covariates to include, each of which can be made after seeing the data
  • Analysing until something works, sometimes called p-hacking, which need not be deliberate and which reliably produces a publishable result from noise
  • Hypothesising after results are known, presenting an exploratory finding as if it had been predicted, which removes the protection that prediction provides
  • Publication bias, since journals preferred novel positive findings and rarely published null results or replications, so the literature shows the successes and hides the failures
  • Incentives, since careers depend on publication counts and prestigious placements rather than on whether findings hold
  • The statistical threshold itself, since a fixed significance cut-off treats a continuous measure of evidence as a binary verdict and encourages exactly the behaviour above

What has changed

The reforms adopted are specific and many are now standard in the affected fields. Preregistration requires researchers to publish their hypothesis, sample size and analysis plan before collecting data, which removes the flexibility that produces false positives, and registered reports go further by having journals accept a study on the basis of its design before results exist, which eliminates publication bias by construction. Power analysis before data collection and substantially larger samples have become expected. Data and code sharing allows others to check analyses directly. Multi-laboratory collaborations run the same protocol across many sites, producing estimates far more reliable than any single study. Journals have created formats for null results and for replications, which were previously close to unpublishable. Some fields have lowered the significance threshold or moved towards reporting effect sizes and intervals rather than binary verdicts. Early evidence suggests preregistered studies report positive results far less often, which is what the reform was supposed to do.

What it does and does not mean

Two misreadings are common and both are worth resisting. The first treats the crisis as evidence that science cannot be trusted, which does not follow, since the crisis was discovered and quantified by scientists using scientific methods and has produced substantial correction, which is the process working rather than failing. The second treats it as confined to psychology, when replication problems have been documented across biomedicine, economics, ecology and beyond, and the underlying incentives are shared. The more useful conclusions are practical: a single study, particularly a small and surprising one, is weak evidence and should be read as such; effect sizes in early reports are systematically inflated; and the fields with the strongest reforms are now producing more trustworthy findings than those without. It also implies something about reading science journalism, since the results most likely to be reported are the novel and surprising ones, which are precisely the ones least likely to replicate.

The takeaway

Large coordinated projects from 2011 onward reproduced well under half of published findings in several fields, with surviving effects around half the original size. The causes were ordinary practice rather than fraud: small samples, flexible analysis choices made after seeing data, presenting exploratory results as predictions, and journals publishing only positive novel findings. Reforms include preregistration, registered reports, data sharing and multi-site collaborations, and preregistered studies report far fewer positive results.

Practise this

Questions from Research Methods and Statistics

Reading about something is not the same as being able to recall it. These are real questions from the Research Methods and Statistics unit in our Science track, answers and explanations included. The unit has 127 in total across 21 steps.

  • Choose all that applyLevel 4

    1. Which are defining features of a true experiment? Pick all that apply.

    • Manipulation of an independent variablecorrect
    • Random assignment to conditionscorrect
    • A comparison or control conditioncorrect
    • Measuring pre-existing groups as they naturally occur

    True experiments manipulate a variable, randomly assign, and compare conditions; measuring natural groups is correlational.

  • Multiple choiceLevel 4

    2. What is the primary purpose of peer review?

    • To have independent experts evaluate a study's quality before publicationcorrect
    • To guarantee the findings are true
    • To repeat the experiment for the authors
    • To calculate the study's p-value

    Peer reviewers check methods, analysis, and reasoning to filter and improve work before it is published.

  • Fact or fibLevel 4

    3. Including a placebo control group lets researchers separate a treatment's specific effect from the improvement people expect just from being treated.

    Answer: True

    A placebo group captures the expectation effect so the treatment's true effect can be isolated.