How Sure Are You of That Average? It Depends How Many You Measured
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
An average calculated from a sample would come out differently with a different sample, and this figure says how much. Confusing it with the spread of the data is extremely common.
What it measures
Taking a sample and calculating its average gives one number, and taking a different sample from the same population would give a slightly different one. Repeating that many times would produce a collection of averages scattered around the true value, and this quantity describes how widely those averages would scatter. It therefore says how precisely the sample average pins down the real one, which is a statement about confidence in a result rather than about the data itself. It falls as the sample grows, because a larger sample gives an average that moves around less.
How it differs from the spread
The two quantities are routinely confused and describe different things:
- •The spread describes how much individual observations vary
- •This describes how much the average would vary between samples
- •The spread does not shrink as more data is collected
- •This does, in proportion to the square root of the sample size
- •Quoting the spread describes the population
- •Quoting this describes confidence in the estimate
Why the square root matters
The relationship with sample size has a practical consequence that governs how studies are designed. Because the figure falls with the square root of the number measured, halving it requires four times the sample, and improving precision by a factor of ten requires a hundred times. That is a harsh return on effort and is why sample sizes are calculated in advance rather than chosen by feel, and why a study that is too small cannot be rescued by analysis. It also means that beyond a point, further measurement buys very little, and effort is better spent on reducing the underlying variability or on measuring something else.
What it assumes
The calculation rests on conditions that are frequently not checked and that invalidate it when they fail. It assumes the observations are independent of each other, so measurements that are clustered, repeated on the same subjects or collected close together in time violate it and the figure comes out too small, overstating confidence. It assumes the sample was drawn at random from the population being described, so a convenience sample gives a precise estimate of the wrong thing. And it describes only random variation, so it says nothing whatever about a systematic error in the instrument or the method, which no amount of data will reduce.
Why graphs mislead with it
Error bars on a chart can show either quantity and frequently do not say which, which makes them genuinely hard to read. Bars showing the spread of the data look large and bars showing confidence in the average look small, so the same data can be drawn to appear noisy or tight depending on an unstated choice. A further trap is that two sets of bars overlapping does not establish that the difference is unimportant, and two sets not overlapping does not establish that it matters, since the relationship between overlap and statistical significance depends on which quantity is drawn. Any chart with error bars should state what they represent, and many do not.
The takeaway
The figure describes how much a sample average would move around if the sample were taken again, so it states confidence in a result rather than the variability of the data, and it falls with the square root of the sample size. Halving it needs four times the measurements. Error bars on charts can show either quantity, and overlap between them settles nothing unless the chart says which is drawn.