← All articles
sciencestatisticsmethodevidenceSeptember 17, 20264 min read

What Is a Confounder? The Third Thing Explaining an Apparent Link

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

Two things can be associated because one causes the other, or because something else causes both. That third factor is a confounder, and identifying and handling them is most of what separates a useful observational study from a misleading one.

What makes something a confounder

A confounder is a factor associated with the supposed cause and independently affecting the outcome, without lying on the causal path between them. The classic illustration is the association between carrying matches and lung cancer, where smoking causes both and accounts for the link entirely. The three requirements matter individually, since a factor associated with the exposure but not affecting the outcome does no damage, one affecting the outcome but unrelated to the exposure adds noise rather than bias, and one lying on the causal path between exposure and outcome is a mediator, and adjusting for it removes part of the real effect rather than a spurious one. That last distinction is the source of a great many analytical errors, since the statistical appearance is identical and only knowledge of the subject can tell them apart.

How they are handled

Several methods address the problem and each has requirements:

  • Randomisation, which distributes all factors evenly whether known or not and is the only method addressing unmeasured ones
  • Restriction, studying only people who share the confounding characteristic, which removes its influence and narrows the applicable population
  • Matching, pairing participants with similar values of the confounder
  • Stratification, analysing within levels of the confounder separately
  • Statistical adjustment in a model, which is the commonest approach and works only for factors that were measured and measured well
  • Instrumental variables and natural experiments, which exploit a source of variation unrelated to the confounders

The ones nobody measured

The fundamental limitation of observational research is that adjustment only works for confounders that were identified and measured, and there is no way to adjust for a factor nobody thought of. That is why an association in observational data supports a causal conclusion far more weakly than the same association in a randomised trial, and the history of medicine contains several cases where large well-conducted observational studies reported benefits that randomised trials subsequently contradicted, with hormone therapy the most cited example. The usual explanation in such cases is that people receiving the treatment differed systematically from those not receiving it in ways connected to health, which is healthy user bias and is difficult to measure. Sensitivity analysis can estimate how strong an unmeasured confounder would have to be to explain an observed association, which is a useful partial answer.

Famous confounded findings

Several widely reported associations turned out to be explained by a third factor, and the examples are worth having. Moderate drinking appeared to be associated with better health outcomes in many observational studies, and a substantial part of that pattern is attributed to the comparison group including people who had stopped drinking because they were already unwell, which is confounding by prior health. Coffee was associated with disease for decades before the association was traced substantially to smoking, since smokers drank more coffee. Studies of vitamin supplements repeatedly found benefits that randomised trials did not reproduce, with the difference attributed to people who take supplements differing systematically in other health behaviours. In each case the original studies were competently conducted and the confounder was either unmeasured or handled inadequately, which is the point rather than a criticism.

Adjusting for the wrong thing

Statistical adjustment is not automatically improving and can introduce bias that was not there. Adjusting for a mediator removes part of the genuine effect, which understates the result. Adjusting for a common consequence of the exposure and the outcome creates a spurious association where none existed, which is a counterintuitive result that follows from the arithmetic and is a known source of error. Adjusting for a variable measured with substantial error leaves residual confounding while giving the appearance of having handled it. And adjusting for many variables chosen because they improve the model fit, rather than because the subject matter implies they are confounders, produces results that cannot be interpreted causally at all. The current recommendation is to decide what to adjust for from an explicit diagram of assumed causal relationships before looking at the data.

The takeaway

A confounder is associated with the exposure, independently affects the outcome and does not lie on the path between them, and only subject knowledge distinguishes it from a mediator. Adjustment handles only measured factors, which is why observational associations support causal claims weakly. Adjusting for a mediator removes real effect and adjusting for a common consequence creates a spurious link.

Practise this

Questions from Research Methods and Statistics

Reading about something is not the same as being able to recall it. These are real questions from the Research Methods and Statistics unit in our Science track, answers and explanations included. The unit has 127 in total across 21 steps.

  • Guess the numberLevel 4

    1. If you run 20 independent significance tests on data with no real effect, using alpha = 0.05, about how many false positives would you expect on average?

    Answer: 1 tests

    With alpha = 0.05, about 5 percent of the 20 tests, or roughly 1, will be a false positive by chance.

  • Choose all that applyLevel 4

    2. Which of these help make a sample more representative of its population? Pick all that apply.

    • Randomly selecting participantscorrect
    • Using a large enough samplecorrect
    • Building a complete sampling framecorrect
    • Recruiting only self-selected volunteers

    Random selection, adequate size, and a good sampling frame reduce bias; relying on volunteers adds self-selection bias.

  • Match the pairsLevel 4

    3. Match each correlation coefficient to its description.

    Answer: r = +0.90 = Strong positive; r = -0.85 = Strong negative; r = +0.15 = Weak positive; r = 0.00 = No linear relationship

    The sign shows direction and the magnitude from 0 to 1 shows the strength of the linear relationship.