← All articles
sciencemethodreasoningevidenceSeptember 17, 20264 min read

What Is a Hypothesis? A Claim Specific Enough to Be Wrong

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

A hypothesis is a proposed answer stated precisely enough that evidence could count against it. That last requirement does most of the work, since a claim compatible with every possible observation is not a weak hypothesis but not one at all.

What makes one usable

A usable hypothesis states something definite about how the world is and implies observable consequences that would not follow if it were false. That second part is what allows it to be tested, since a test consists of checking whether the implied consequence occurs. Specificity matters, because a vague claim can accommodate any result and therefore learns nothing from any test, and precision is what exposes a claim to failure. The requirement is frequently summarised as falsifiability, and the summary is useful and slightly too simple, since no hypothesis is tested in isolation and a failed prediction can always be blamed on an auxiliary assumption rather than on the hypothesis itself. That complication is real and does not remove the basic point, which is that a claim with no possible disconfirming observation is doing something other than describing the world.

Where they come from

Generating one is a different activity from testing one and is far less well described:

  • Observation of a pattern that seems to need explaining, which is the textbook route and is not the commonest
  • Analogy with a mechanism known to operate elsewhere, which supplies most working hypotheses in practice
  • Extension of an existing theory to a new case, where the hypothesis is what the theory predicts
  • Anomaly, where something fails to fit and the hypothesis is an account of why
  • Systematic elimination of alternatives, which narrows towards a candidate
  • Accident, which accounts for more than the official picture admits, since the discovery is generally reconstructed afterwards as though it followed a method

The statistical version

In quantitative work the term takes on a specific technical meaning that differs from the general one and causes confusion. The procedure states a null hypothesis, typically that there is no effect or no difference, and an alternative, and then asks how probable the observed data would be if the null were true. A result unlikely enough under the null leads to rejecting it. That is a narrower operation than it appears, since rejecting the null does not establish the alternative, since the probability calculated is of the data given the hypothesis rather than the reverse, and since the conventional threshold is arbitrary. Misreading the output as the probability that the hypothesis is true is the single most common error in reading quantitative research, and it is made regularly by people who should know better, including in published papers.

Competing explanations

A single hypothesis tested in isolation is a weaker arrangement than several tested against each other, and the reasoning is worth stating. Testing one explanation asks only whether the data are consistent with it, and data consistent with one explanation are frequently consistent with several, so a confirming result discriminates very little. Designing a study so that rival explanations predict different outcomes converts the same effort into a result that eliminates something regardless of which way it goes. The approach has a long history under the name of strong inference and is standard in some fields and neglected in others. It also addresses the attachment problem structurally, since a researcher holding several candidates has less invested in any one. The practical question to ask of any study is what result would have been reported had the favoured explanation been wrong, and whether the design could have produced it.

Living with the wrong ones

Most hypotheses turn out to be wrong, which is the normal condition of research rather than a failure, and handling that well is a practical skill. Attachment to a hypothesis is the recognised occupational hazard, since a researcher who has invested years in an idea has reasons beyond evidence to want it to hold, and the standard protections are procedural rather than moral, including preregistering what will count as a test, inviting criticism early and designing studies that could return a clear negative. Holding several competing hypotheses at once is advised by many researchers on the grounds that it is harder to become attached to one of several, and the approach of designing experiments to discriminate between rival explanations rather than to confirm a favoured one is considerably more efficient. A wrong hypothesis clearly tested is a contribution, and a vague one never tested is not.

The takeaway

A hypothesis must imply observations that would not occur if it were false, which is what makes it testable, and a claim compatible with every result is not a weak one but not a hypothesis. The statistical version asks how probable the data are given no effect, which is not the probability that the hypothesis is true. Most turn out wrong, and the protections against attachment are procedural.

Practise this

Questions from The Scientific Method

Reading about something is not the same as being able to recall it. These are real questions from the The Scientific Method unit in our Science track, answers and explanations included. The unit has 131 in total across 22 steps.

  • Build the sentenceLevel 2

    1. Build a sentence about how science begins.

    Answer: Scientists start with a question they can test

    Science usually begins with a clear question that can actually be tested.

  • Match the pairsLevel 2

    2. Match each thing you want to measure to the best tool for the job.

    Answer: How long a reaction takes = Stopwatch; The mass of a sample = Balance; The temperature of water = Thermometer; The volume of a liquid = Measuring cylinder

    Choosing the correct tool gives accurate, useful measurements.

  • Match the pairsLevel 2

    3. Match each result to what it tells you.

    Answer: Data matches the prediction = Hypothesis supported; Data does not match the prediction = Hypothesis not supported; Results are similar every trial = Data is reliable; Results vary wildly each trial = Data is not reliable

    Matching results support a hypothesis, and consistent trials show the data is reliable.