How Do You Check a Laboratory Is Any Good? Send Everyone the Same Sample
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
Distributing identical samples to many laboratories and comparing the answers is the only practical way to find out whether their results actually mean anything.
How the scheme works
An organising body prepares a large quantity of a material, confirms it is uniform, divides it into portions and sends one to every participating laboratory without telling them what the answer is. Each laboratory analyses it by its normal procedure and reports a result. The organiser then compares all the results, establishes a reference value, and tells each participant how far from that value they were and how they compare with everybody else. The exercise is repeated regularly, so a laboratory sees its performance over time rather than on a single occasion.
What it detects
The comparison catches faults that internal checks cannot:
- •A consistent bias in one direction, invisible from inside
- •A miscalibrated instrument that repeats itself precisely
- •A procedure being followed differently than intended
- •A calculation or unit error in reporting
- •Contamination affecting every sample the same way
- •Drift over time, seen by comparing successive rounds
Why internal checks are not enough
A laboratory can be highly repeatable and consistently wrong, and nothing available inside the building will reveal it. Running the same sample ten times measures precision, which is agreement with yourself, and says nothing about accuracy, which is agreement with the truth. A wrongly calibrated instrument gives the same wrong answer every time and looks excellent by any internal measure. Purchased reference materials help, since they carry a certified value, and they are expensive and limited in range. Comparison against many independent laboratories is the broadest available check, because a shared error is unlikely when methods differ.
How the reference value is set
Deciding what the right answer was is the delicate part of running such a scheme, since the organiser frequently does not know it either. Where the material can be prepared by adding a known quantity to a clean base, the value is known by construction and is the strongest option. Where it cannot, the value is taken from a small number of expert laboratories using definitive methods, or from a robust average of all participants that discounts extreme results. The last option has an obvious weakness, namely that a shared systematic error across the whole field would be invisible, which is why independent methods matter.
Where it matters
Participation is not optional in most regulated fields and the consequences of failure are real. Clinical laboratories reporting patient results must participate to keep accreditation, and repeated poor performance can close a service. Forensic laboratories are subject to the same requirement, and failures have been used in court to challenge evidence. Environmental monitoring, food safety, water testing and materials certification all operate under such schemes. A less formal version appears in research, where an unexplained disagreement between groups measuring the same quantity is usually resolved by circulating identical samples.
The takeaway
Sending identical portions of one material to many laboratories and comparing answers exposes a consistent bias that no internal check can find, because repeating a measurement tests agreement with yourself rather than with the truth. Accreditation in clinical, forensic and environmental testing depends on taking part, and repeated failure can close a service.