Why Ask a Thousand People About Millions? Sampling That Actually Works
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
A well-drawn sample of a thousand estimates a population of millions to within a few percentage points, which sounds implausible and is correct. The size of the population barely matters and how the sample was drawn matters enormously.
Why a small sample works
The accuracy of an estimate from a sample depends on the size of the sample and almost not at all on the size of the population it came from, which is the counterintuitive result at the centre of the whole subject. The reason is that the variability of a sample average falls as the square root of the sample size, and the population size enters only through a correction that is negligible unless the sample is a substantial fraction of the whole. That is why a national poll and a city poll need similar numbers for similar accuracy, and it is why increasing a sample from a thousand to four thousand only halves the error, which is a poor return and is why polls are the size they are.
What the margin of error covers
The quoted figure is narrower in meaning than most readers assume:
- •It describes sampling variability alone, meaning the luck of who happened to be selected
- •It assumes the sample was drawn at random from the population of interest
- •It says nothing about people who could not be reached or refused to answer
- •It says nothing about whether the question was well worded or understood
- •It says nothing about whether people answered honestly or knew their own mind
- •Those other sources of error are frequently larger than the quoted margin
Where sampling goes wrong
The failures are about who ends up in the sample rather than about how many. A sample drawn from a list that excludes part of the population cannot represent it, which is how a famous 1936 poll predicted the wrong result from over two million responses, having drawn its names from telephone directories and vehicle registrations during a depression. Low response rates are the modern version, since the small proportion who answer may differ systematically from those who do not, and response rates to telephone polls have fallen into the single figures. Self-selected samples, including online polls and volunteer panels, represent whoever chose to participate. Weighting adjusts for known differences and cannot correct for unknown ones, which is where recent polling failures have been located.
What weighting does and does not fix
Every real survey is adjusted before publication and understanding that adjustment matters. Weighting compares the sample against known population figures on characteristics including age, sex, region and education, and gives more weight to respondents from under-represented groups so the weighted sample matches the population on those variables. That corrects for differences in who responded on the characteristics used. It cannot correct for differences on anything not used, and the recent failures in political polling have been attributed largely to differences in willingness to respond that correlate with political views themselves, which no demographic weighting reaches. Heavy weighting also increases variability, since a few respondents end up carrying a great deal of the estimate, which widens the real error beyond the quoted margin.
How proper samples are drawn
Several designs exist and each addresses a practical obstacle. Simple random sampling gives every member an equal chance and requires a complete list of the population, which frequently does not exist. Stratified sampling divides the population into groups and samples within each, which guarantees representation of small groups and improves precision where the groups differ. Cluster sampling selects whole groups such as households or districts and surveys within them, which is far cheaper for face-to-face work and reduces precision for a given size. Systematic sampling takes every nth member from a list. Multi-stage designs combine these. Whatever the design, the essential requirement is that selection is determined by a known probability rather than by convenience or by the respondent.
The takeaway
Accuracy depends on sample size and barely on population size, and error falls as the square root of the sample, so quadrupling it only halves the error. The quoted margin covers sampling luck alone and says nothing about non-response, question wording or honesty, which are frequently larger. A 1936 poll with over two million responses got the result wrong because of who was on its lists.