← All articles
scienceresearchmethodplanningSeptember 17, 20263 min read

Why Run a Small Version First? Finding Out What Will Go Wrong

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

A small preliminary version of a study exists to test whether the procedures work, not to find out whether the hypothesis is true. Treating it as evidence about the answer is a common and serious error.

What it is for

The purpose is to test the machinery of a study rather than its question. Recruitment may be slower than expected, the questionnaire may be misunderstood, the equipment may fail, the task may take twice as long as planned, participants may drop out in numbers nobody anticipated, and the data may arrive in a form that cannot be analysed as intended. Every one of those is far cheaper to discover in a small version than in a full one, and each has ruined studies that went straight to full scale. Running a reduced version first therefore protects the investment in the main study, which is a question about feasibility rather than about findings.

What it should report

The outcomes that matter are practical:

  • How many people were approached and how many agreed
  • How long each stage actually took
  • How many participants completed and how many left
  • Whether the measures were understood as intended
  • What went wrong with equipment, procedures or data handling
  • Whether the estimated cost and timetable were remotely accurate

Why the results should not be reported as findings

A small study has a low probability of detecting a real effect and a high probability of producing a misleading estimate, which is the central statistical reason for not treating its outcome as evidence. An effect that reaches significance in a very small sample is almost certainly overestimated, because only a large apparent effect could reach significance at that size, which means the figure published is systematically inflated and any later study powered from it will be underpowered. Publishing a promising preliminary result also creates pressure not to contradict it. The correct output is a statement about feasibility, plus the variability observed, which is what a proper calculation of the required sample size actually needs.

What a proper one leads to

The output feeds directly into designing the full study, and knowing how makes the exercise worthwhile rather than a formality. Observed recruitment rates give a realistic timetable and indicate how many sites are needed. Observed dropout determines how many participants must be recruited to end with the required number. The variability seen in the outcome measure is the figure a sample size calculation needs, and guessing it is the commonest reason full studies end up underpowered. Problems with procedures produce concrete changes to the protocol, which should be documented rather than quietly made. And an honest pilot sometimes concludes that the full study is not feasible as designed, which is a useful result and is published far less often than it occurs.

How it gets misused

The label is attached to studies that are not pilots at all, and the pattern is recognisable. A small study that found nothing is described as preliminary, which implies the question remains open when the honest report is a null result. A small study that found something is published as a finding, with the preliminary status mentioned once and forgotten. Repeated pilots of the same intervention accumulate without any full study ever being conducted, which is common in health services research and produces a literature full of promising beginnings and no conclusions. Funders and journals have responded with registration requirements and with explicit guidance that feasibility studies report feasibility outcomes rather than effectiveness.

The takeaway

A preliminary version tests recruitment, timing, dropout, comprehension and data handling, all of which are far cheaper to fix at small scale. Its outcome is a statement about feasibility plus the variability observed, which is what a sample size calculation needs. An effect reaching significance in a very small sample is systematically inflated, so reporting it as a finding misleads and produces underpowered follow-up studies.

Practise this

Questions from Research Methods and Statistics

Reading about something is not the same as being able to recall it. These are real questions from the Research Methods and Statistics unit in our Science track, answers and explanations included. The unit has 127 in total across 21 steps.

  • Fact or fibLevel 4

    1. A negative correlation means that as one variable increases, the other tends to decrease.

    Answer: True

    Negative (inverse) correlations have a downward trend: higher values of one go with lower values of the other.

  • Build the sentenceLevel 4

    2. Build a sentence describing what peer review does.

    Answer: Peer review lets experts check a study

    Peer review is independent experts checking a study's quality before it is published.

  • Match the pairsLevel 4

    3. Match each significance-testing term to its meaning.

    Answer: Type I error = Rejecting a true null (false positive); Type II error = Failing to reject a false null (false negative); Alpha = Accepted Type I error rate; Statistical power = Chance of detecting a real effect

    Type I and Type II errors, alpha, and power are the core trade-offs of hypothesis testing.