← All articles
sciencemethodexperimentsevidenceSeptember 17, 20264 min read

What Is a Control Group? The Comparison That Makes a Result Mean Something

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

Observing that people improved after a treatment establishes almost nothing, because people improve for many reasons. A control group supplies what would have happened otherwise, and without that comparison there is no way to attribute the change to anything.

The problem it solves

Any measurement taken before and after an intervention will differ, and the difference has many possible causes that have nothing to do with the intervention. Conditions improve on their own, since many illnesses resolve and many problems fluctuate. Extreme measurements tend to be followed by less extreme ones simply because extremes are partly due to chance, which is regression to the mean and which alone produces apparent improvement in anyone selected for being unwell. Expectation changes outcomes and reported outcomes, which is the placebo effect and its reporting counterpart. Being observed changes behaviour. Other things happen during the study period. A control group receiving everything except the intervention experiences all of these and none of the treatment, so the difference between the groups isolates what the treatment did, which nothing about the treated group alone can establish.

How the comparison is protected

The comparison only works if the groups differ in the intervention alone, which requires deliberate machinery:

  • Randomisation, which assigns participants by chance so that known and unknown differences are distributed evenly between groups
  • Blinding of participants, so expectation operates equally on both sides
  • Blinding of those assessing outcomes, since judgement of an outcome is influenced by knowing the assignment
  • A placebo or sham intervention, so the control experiences the same procedure without the active component
  • Allocation concealment, so whoever enrols participants cannot predict the next assignment and steer it
  • Analysing everyone in the group they were assigned to, regardless of what they actually received, which preserves the randomisation

Where the idea came from

The design took a long time to become standard and the history explains why its features exist. An early and often cited trial compared treatments for scurvy aboard a ship in the eighteenth century, assigning different remedies to pairs of sailors, which contains the core idea of a simultaneous comparison. Agricultural research in the early twentieth century contributed the statistical foundations, since comparing crop treatments across variable fields required a formal theory of how to assign plots, and randomisation entered the literature there before it entered medicine. The first widely recognised randomised medical trial tested a tuberculosis treatment in the 1940s, with allocation concealed and assessment blinded, and it became the template. Regulatory requirements followed a series of drug disasters, and the modern expectation that a new treatment must demonstrate benefit against a control before approval is a twentieth-century development rather than an ancient principle.

When you cannot have one

A great deal of important research cannot use this design and the alternatives are weaker in known ways. Withholding a treatment believed to work is unethical, so new treatments are usually compared against the current standard rather than against nothing. Some exposures cannot be assigned, since nobody can randomise people to smoke or to experience a disaster, so the evidence comes from observation and carries the permanent possibility that the groups differed beforehand in ways that explain the outcome. Natural experiments exploit situations where something approximating random assignment occurred by accident, including policy changes affecting arbitrary groups. Statistical adjustment attempts to compensate for measured differences and cannot address unmeasured ones. Historical comparison uses earlier patients as the control and is unreliable because many things change over time. Each of these is a compromise, and knowing which compromise was made is most of what reading a study well consists of.

How the comparison gets broken

Even a properly designed study can lose the protection the control provides. Dropout that differs between groups reintroduces the differences randomisation removed, since participants leaving because a treatment is unpleasant or ineffective changes who remains. Unblinding happens when a treatment has obvious effects, and participants and assessors frequently guess correctly, which restores the expectation effects blinding was meant to remove. Contamination occurs when control participants obtain the treatment independently. Small groups leave randomisation unable to balance characteristics reliably, which is why small trials produce unstable results. Changing outcomes after seeing the data, or reporting only some of those measured, breaks the comparison at the analysis stage rather than the design stage. Each of these is checkable in a published report, and the presence of a control group is the beginning of an assessment rather than the end of one.

The takeaway

Measurements change for many reasons including natural resolution, regression to the mean, expectation and being observed, so a before and after comparison attributes all of that to the treatment. A control group experiences everything except the intervention, and the difference isolates the effect. Randomisation, blinding and analysing people in their assigned groups protect the comparison, and dropout and unblinding break it.

Practise this

Questions from Data, Graphs and Evidence

Reading about something is not the same as being able to recall it. These are real questions from the Data, Graphs and Evidence unit in our Science track, answers and explanations included. The unit has 131 in total across 22 steps.

  • Fact or fibLevel 3

    1. The dependent variable (the one you measure) is usually plotted on the vertical y-axis.

    Answer: True

    By convention the measured dependent variable goes on the y-axis and the variable you control on the x-axis.

  • Match the pairsLevel 4

    2. Match each measure to its definition.

    Answer: Mean = Sum divided by how many values; Median = Middle value when ordered; Mode = Most common value; Range = Highest minus lowest

    Mean, median, mode, and range each summarise a data set in a different way.

  • Fill the blankLevel 3

    3. Results are described as ____ if repeating the same experiment gives closely similar values each time.

    • reliablecorrect
    • biased
    • random
    • unfair

    Reliable (repeatable) results are ones you get again when the experiment is repeated.