← All articles
mathstatisticsevidenceresearchSeptember 17, 20264 min read

What Is a Sample Size? Why Small Studies Mislead in Both Directions

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

How many people or measurements a study includes determines what it can show, and the relationship is not intuitive. Small samples do not simply give less certain answers, they give answers that are wrong in characteristic ways, including producing effects that look larger than the truth.

Why bigger narrows the range

The variability of a sample average falls as the sample grows, and it does so in proportion to the square root of the number of observations, which has an important practical consequence: halving the width of a confidence interval requires four times as many observations, not twice as many. That square root relationship governs the economics of research, since the cost of precision rises steeply and there is a point beyond which additional data buys very little. It also explains why small samples produce wildly variable results, since an average of ten observations can land far from the true value by chance while an average of a thousand rarely does. A related and frequently misunderstood point is that what matters is the absolute size of the sample rather than the fraction of the population it represents, so a properly drawn sample of a thousand describes a country of sixty million about as well as it describes a town of sixty thousand, which is counterintuitive and is why national polls use the sample sizes they do.

How small samples mislead

The failures are systematic rather than merely random, and several are well documented:

  • Low power, meaning a genuine effect is likely to be missed, so a negative result from a small study is weak evidence of absence
  • Effect size inflation, since in an underpowered study only the largest chance deviations reach significance, so any effect that is published is exaggerated, which is called the winner's curse
  • Reduced probability that a significant result is true, because a low-powered study produces a higher ratio of false positives to true positives among its significant findings
  • Extreme values in small units, which is why the smallest schools, hospitals and regions appear at both the top and the bottom of performance tables, a pattern misread as evidence that small institutions are better
  • Instability, since adding a few observations can change a small study's conclusion entirely
  • Subgroup analysis, where dividing an already modest sample produces groups small enough that apparent differences between them are mostly noise

How researchers choose a number

A formal power calculation is performed before data collection and combines four quantities: the smallest effect considered worth detecting, the variability expected in the measurements, the significance threshold to be used, and the desired probability of detecting the effect if it exists, conventionally set at eighty or ninety percent. Fixing three of those determines the fourth, so specifying a target effect and an acceptable miss rate yields the required sample. The weakness of the procedure is that the expected effect size and variability must be estimated in advance, frequently from small preliminary studies that suffer from exactly the inflation described above, so calculations are routinely optimistic and studies routinely end up underpowered. Related practices matter as much as the number itself, including registering the intended sample and analysis in advance so that collection does not stop when the result looks good, which is a documented source of false positives, and reporting confidence intervals rather than only whether a threshold was crossed.

Where size is not the problem

A large sample fixes random error and does nothing about bias, which is the more dangerous failure because it does not announce itself. A biased sampling method produces a confidently wrong answer, and the classic illustration is a 1936 American election poll with over two million responses that predicted the wrong winner because its respondents were drawn from lists of telephone and car owners during a depression, while a much smaller but properly drawn sample got it right. Non-response bias operates the same way, since the people who answer differ systematically from those who do not, and response rates to surveys have fallen dramatically in recent decades, which is a larger threat to polling accuracy than sample size. Measurement error, poorly worded questions and researcher degrees of freedom in analysis are likewise unaffected by collecting more data. The practical rule is that sample size determines precision while sampling method determines accuracy, and a precise wrong answer is worse than an imprecise right one.

The takeaway

Precision improves with the square root of the number of observations, so four times the data is needed to halve the uncertainty, and the absolute size matters rather than the fraction of the population. Underpowered studies both miss real effects and exaggerate the ones they find, since only large chance deviations reach significance. A large sample fixes random error and leaves bias untouched, which is why a two million response poll called an election wrong.

Practise this

The Math track

From counting and shapes to algebra, calculus and beyond - number magic one tiny step at a time.

18 units and 2,161 questions, each with a written explanation. Every unit page shows what it covers and real example questions before you start.