Why Is the Same Bell Shape Everywhere? Adding Things Up
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
Heights, measurement errors, examination marks and a great many other quantities cluster symmetrically around an average in the same characteristic shape. There is a mathematical reason why, and knowing it shows where the shape does not apply.
What the shape describes
The distribution is symmetric about its average, with values close to that average being common and values further away becoming rapidly rarer in both directions. It is described entirely by two numbers, the average and a measure of how spread out the values are, so stating those two fixes the whole shape. A fixed proportion of values falls within one unit of spread either side of the average, around sixty eight per cent, with about ninety five per cent within two units and over ninety nine within three. Those proportions hold regardless of what is being measured, which is what makes the shape so useful and why it is worth recognising.
Why it appears so often
The reason is a theorem rather than a coincidence:
- •Many quantities result from adding together many small independent contributions
- •Height results from many genetic and environmental factors, each contributing a little
- •Measurement error results from many small independent sources of error
- •Adding many independent random contributions produces this shape
- •That holds regardless of the shape of the individual contributions
- •The result is the central limit theorem and is among the most important in the subject
Where it does not apply
Assuming the shape where it does not hold produces serious errors, and the failures are identifiable in advance. Quantities produced by multiplying rather than adding contributions are skewed rather than symmetric, which describes incomes, city sizes and many biological measures. Quantities where one extreme event dominates, including financial returns, earthquake energies and insurance claims, have far heavier tails, meaning extreme values are vastly more common than this shape predicts, and treating them as though they followed it understates risk by enormous factors. Quantities bounded at zero cannot be symmetric if the average is close to that bound. And quantities produced by two different processes mixed together produce two humps rather than one.
Why the same shape keeps reappearing
The theorem behind the shape is stronger and stranger than it first appears, and understanding what it actually says is worthwhile. It states that the average of many independent random quantities approaches this shape as the number grows, whatever distribution those quantities individually follow, provided each contributes a modest share and their spread is finite. That means adding dice rolls, coin flips, waiting times or anything else converges to the same curve, which is why it appears in contexts with no connection to each other. The conditions matter, and each one names a way it can fail, since contributions that are not independent, that are dominated by one large one, or that have infinite spread all break the result.
How it is misused
Several misapplications recur and are worth naming. Grading on a curve forces examination results into the shape whether or not the underlying performance took it, which converts an absolute judgement into a relative one and is defensible only if the cohort is large and comparable. Assuming it in financial risk modelling understated the probability of large movements catastrophically before 2008. Describing human characteristics as normally distributed and then reasoning about the tails is common in arguments about ability and is frequently done with data that does not have the shape. And reporting an average without a spread is meaningless for any distribution and is especially misleading where the shape is skewed.
The takeaway
The shape is symmetric about its average and fixed entirely by that average and a measure of spread, with about sixty eight per cent of values within one unit of spread and ninety five within two. It appears because adding many small independent contributions produces it regardless of their individual shapes. It fails for quantities produced by multiplication, for those dominated by extremes, and for those bounded near zero.