What Is an Effect Size? How Big the Difference Is, Not Whether It Exists
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
A result can be statistically significant and far too small to matter. Effect size measures how large a difference actually is, separately from whether it can be distinguished from chance, and reporting it is the correction for a long-standing bad habit.
The distinction
A significance test asks whether an observed difference is larger than would plausibly arise by chance if there were no real difference, and the answer depends heavily on how many observations were made. A tiny difference becomes statistically detectable with a large enough sample, and a large difference can fail to reach significance in a small one, which means significance conflates the size of the effect with the size of the study. Effect size separates them by reporting the magnitude of the difference in interpretable units, either in the original units of measurement or standardised so that studies using different measures can be compared. The two questions are genuinely different, since one asks whether an effect exists and the other asks whether it is worth anything, and a study answering only the first has left the practically important question untouched.
The common measures
Several standard quantities appear depending on what is being compared:
- •A standardised mean difference, expressing the gap between two groups in units of the variation within them, which is the most widely reported
- •Correlation coefficients, which express how strongly two measures move together
- •Odds and risk ratios, used where the outcome is a yes or no event, and which are frequently misread
- •Absolute risk difference, which answers how many additional people are affected and is usually more informative than a ratio
- •Number needed to treat, the number of patients who must receive a treatment for one to benefit, which is directly interpretable
- •Proportion of variance explained, which describes how much of the variation a factor accounts for
Why ratios mislead
The gap between relative and absolute measures is the source of a great deal of misleading reporting, and it is worth understanding precisely. A treatment that halves a risk sounds impressive, and if the original risk was two in a million the treatment prevents one case per million people, which is a real but negligible benefit. The same relative reduction applied to a risk of forty percent prevents twenty cases per hundred, which is transformative. The ratio is identical and the practical meaning is not remotely comparable. Press coverage and promotional material favour the relative figure because it is larger, and the absolute figure is frequently omitted, which is a known and documented pattern rather than an occasional lapse. The correction is simple, since asking what the risk was beforehand converts any relative claim into an absolute one, and any report that does not supply the baseline is incomplete.
Planning a study around it
Effect size is not only a way of reporting results but the quantity a study has to be designed around, which is where it does the most useful work. Statistical power, the probability that a study will detect an effect if one exists, depends on the size of the effect, the number of observations and the threshold chosen, and calculating the sample size needed requires committing to how large an effect is worth detecting. That forces a useful question at the design stage, since specifying the smallest effect that would matter practically is a substantive judgement about the subject rather than a statistical one. Studies that skip this step are frequently too small to detect anything but implausibly large effects, which wastes the participants and produces a literature of inconclusive results. It also guards against the opposite problem, since a very large study will detect differences too small to be of any use.
Interpreting the numbers
Conventional labels exist for small, medium and large standardised effects, and their originator explicitly described them as rough guidance for cases where nothing better was available. They are now applied mechanically, which has drawbacks, since what counts as a large effect depends entirely on the field and the question, and a small effect on a common outcome can matter far more than a large effect on a rare one. Effects in fields studying human behaviour are typically small by those conventions, which reflects the number of factors influencing any outcome rather than a failure of the research. Confidence intervals around an effect estimate matter more than the point estimate, since they show the range of values the data are consistent with, and a study reporting an effect size without an interval has not said how precisely it was measured. Comparing effect sizes across studies is the basis of meta-analysis and depends on the measures being genuinely comparable.
The takeaway
Significance conflates the size of an effect with the size of the study, since any real difference becomes detectable with enough observations. Effect size reports magnitude separately. A halved risk means almost nothing if the risk was two in a million and a great deal if it was forty percent, which is why relative figures without a baseline are incomplete. The conventional small, medium and large labels were offered as rough guidance.