How Does Opinion Polling Work? Asking a Thousand People About Sixty Million
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
A well-conducted survey of about a thousand people estimates the views of an entire country to within roughly three percentage points, and the size of the country makes almost no difference to that figure. This is genuinely counterintuitive and follows from the mathematics of sampling, and the reason polls nevertheless go wrong is almost never the sample size. It is who agrees to answer, who turns out to vote, and how the answers are adjusted afterwards.
Why a thousand is enough
The precision of an estimate from a random sample depends on the size of the sample and essentially not at all on the size of the population it is drawn from, provided the population is much larger than the sample. The reason is that what the sample measures is the proportion in the population, and a random draw is equally informative whether the pot contains a million or a hundred million. The margin of error falls with the square root of the sample size, which has an awkward consequence: halving the error requires quadrupling the sample and therefore roughly quadrupling the cost, which is why almost every published poll uses between eight hundred and two thousand respondents. The familiar plus or minus three points describes a range that would contain the true value in ninety-five of a hundred repeated samples, and it applies only to the whole sample, so any subgroup reported from the same poll has a much larger error that is rarely printed.
The part that actually goes wrong
Sampling error is the smallest source of trouble. The larger problems are structural:
- •Non-response, since response rates for telephone polling have collapsed from around thirty-five percent in the 1990s to low single digits, and the people who still answer differ systematically from those who do not
- •Coverage, meaning whether the method can reach everyone, which was the problem with landline-only sampling and is now the problem with online panels of volunteers
- •Turnout modelling, since an election poll must estimate not what everyone thinks but what those who actually vote think, and identifying likely voters is a judgement rather than a measurement
- •Weighting, in which results are adjusted to match known population characteristics for age, sex, region, education and past vote, which corrects known imbalances and cannot correct unknown ones
- •Question wording and order, which measurably shift answers, so that asking about a related topic first changes what follows
- •Herding, in which pollsters whose result differs sharply from the average quietly adjust or decline to publish, which makes the published spread narrower than the true uncertainty
The famous failures
Each large polling miss taught a specific lesson. The Literary Digest predicted a landslide for Alf Landon in 1936 from a survey of over two million people and was wrong by an enormous margin, because the sample was drawn from telephone directories and car registrations, which in the Depression skewed wealthy; George Gallup predicted the correct result from a few thousand properly sampled respondents, which established the field. In 1948 the polls stopped weeks before the vote and missed a late shift, producing a newspaper headline announcing the wrong winner. British polls in 1992 and again in 2015 understated Conservative support, and the 2015 inquiry concluded the samples were unrepresentative rather than the respondents dishonest. The 2016 American polls were reasonably accurate nationally and missed several states, largely by failing to weight for education, which had become strongly associated with vote choice for the first time. The recurring pattern is that polls fail when the electorate changes in a way the weighting scheme was not built to capture.
Reading one properly
A few habits make published polling far more useful. Look at the trend across many polls rather than any single one, since averages are consistently more accurate than individual results and aggregation cancels house effects, the systematic tendency of a particular pollster to lean one way. Check the field dates, because a poll taken before an event does not reflect it. Check the sample and the method, since an online panel, a telephone poll and a text-message survey have different biases. Treat subgroup numbers with suspicion. Note that a lead within the margin of error is not a tie, since the probability of one candidate leading is higher than the binary framing suggests, and that the margin applies to each figure separately so the error on a difference is larger. And remember what a poll is: a measurement of opinion at a moment, not a forecast, and the conversion of polling into a probability of winning involves modelling assumptions that are separate from the polling itself.
The takeaway
A random sample of about a thousand estimates a population proportion to roughly three points regardless of how large the population is, because precision depends on sample size rather than on the size of the pot. The errors that matter come from non-response now in low single digits, from coverage, from guessing who will actually vote, and from weighting that can correct known imbalances and not unknown ones. Historical misses in 1936, 1992, 2015 and 2016 each traced to an unrepresentative sample or a missing adjustment, and averaging many polls beats trusting one.