What Is a Hypothesis? A Claim Specific Enough to Be Wrong
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
A hypothesis is a proposed answer stated precisely enough that evidence could count against it. That last requirement does most of the work, since a claim compatible with every possible observation is not a weak hypothesis but not one at all.
What makes one usable
A usable hypothesis states something definite about how the world is and implies observable consequences that would not follow if it were false. That second part is what allows it to be tested, since a test consists of checking whether the implied consequence occurs. Specificity matters, because a vague claim can accommodate any result and therefore learns nothing from any test, and precision is what exposes a claim to failure. The requirement is frequently summarised as falsifiability, and the summary is useful and slightly too simple, since no hypothesis is tested in isolation and a failed prediction can always be blamed on an auxiliary assumption rather than on the hypothesis itself. That complication is real and does not remove the basic point, which is that a claim with no possible disconfirming observation is doing something other than describing the world.
Where they come from
Generating one is a different activity from testing one and is far less well described:
- •Observation of a pattern that seems to need explaining, which is the textbook route and is not the commonest
- •Analogy with a mechanism known to operate elsewhere, which supplies most working hypotheses in practice
- •Extension of an existing theory to a new case, where the hypothesis is what the theory predicts
- •Anomaly, where something fails to fit and the hypothesis is an account of why
- •Systematic elimination of alternatives, which narrows towards a candidate
- •Accident, which accounts for more than the official picture admits, since the discovery is generally reconstructed afterwards as though it followed a method
The statistical version
In quantitative work the term takes on a specific technical meaning that differs from the general one and causes confusion. The procedure states a null hypothesis, typically that there is no effect or no difference, and an alternative, and then asks how probable the observed data would be if the null were true. A result unlikely enough under the null leads to rejecting it. That is a narrower operation than it appears, since rejecting the null does not establish the alternative, since the probability calculated is of the data given the hypothesis rather than the reverse, and since the conventional threshold is arbitrary. Misreading the output as the probability that the hypothesis is true is the single most common error in reading quantitative research, and it is made regularly by people who should know better, including in published papers.
Competing explanations
A single hypothesis tested in isolation is a weaker arrangement than several tested against each other, and the reasoning is worth stating. Testing one explanation asks only whether the data are consistent with it, and data consistent with one explanation are frequently consistent with several, so a confirming result discriminates very little. Designing a study so that rival explanations predict different outcomes converts the same effort into a result that eliminates something regardless of which way it goes. The approach has a long history under the name of strong inference and is standard in some fields and neglected in others. It also addresses the attachment problem structurally, since a researcher holding several candidates has less invested in any one. The practical question to ask of any study is what result would have been reported had the favoured explanation been wrong, and whether the design could have produced it.
Living with the wrong ones
Most hypotheses turn out to be wrong, which is the normal condition of research rather than a failure, and handling that well is a practical skill. Attachment to a hypothesis is the recognised occupational hazard, since a researcher who has invested years in an idea has reasons beyond evidence to want it to hold, and the standard protections are procedural rather than moral, including preregistering what will count as a test, inviting criticism early and designing studies that could return a clear negative. Holding several competing hypotheses at once is advised by many researchers on the grounds that it is harder to become attached to one of several, and the approach of designing experiments to discriminate between rival explanations rather than to confirm a favoured one is considerably more efficient. A wrong hypothesis clearly tested is a contribution, and a vague one never tested is not.
The takeaway
A hypothesis must imply observations that would not occur if it were false, which is what makes it testable, and a claim compatible with every result is not a weak one but not a hypothesis. The statistical version asks how probable the data are given no effect, which is not the probability that the hypothesis is true. Most turn out wrong, and the protections against attachment are procedural.