What Is DNA Profiling? Counting Repeats, Not Reading Genes
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
DNA profiling does not sequence a person's genome or read anything about their health or appearance. It counts how many times a short sequence repeats at each of a set of specific locations, chosen precisely because they vary enormously between people and carry no known functional information, which is what makes the technique both powerful and comparatively unrevealing.
What is measured
Scattered through the genome are short tandem repeats, places where a short motif of a few bases repeats consecutively, and the number of repeats at a given location varies widely between individuals. Standard forensic profiling examines a set of such locations, currently around twenty in the main national databases, plus a marker indicating sex. At each location a person has two values, one inherited from each parent, and the combination across all locations is the profile. The probability of two unrelated people sharing a full profile is astronomically small, commonly quoted as one in many billions, which is what gives the evidence its weight. Crucially, these locations were selected to sit outside genes and to have no known association with any trait, so a standard profile reveals nothing about health, ancestry or appearance beyond sex, which was a deliberate design choice made to limit the privacy intrusion of holding the data.
How a sample becomes a profile
The laboratory process is standardised and each stage introduces its own constraints:
- •Recovery from a sample, which may be blood, semen, saliva, hair roots or skin cells transferred by touch
- •Extraction and quantification of DNA, which determines whether there is enough to proceed and whether it is degraded
- •Amplification by polymerase chain reaction, which copies the target regions billions of times and is what makes work from tiny samples possible
- •Separation by size using capillary electrophoresis, producing a chart of peaks whose positions give the repeat numbers
- •Interpretation, which is straightforward for a single-source sample of good quality and becomes genuinely difficult for mixtures and degraded material
- •Comparison against a suspect's reference sample or a search of a national database, followed by a statistical statement of how rare the profile is
Where the difficulties are
The technique's reputation for certainty applies to a clean single-source sample and degrades sharply outside that. Mixtures containing DNA from several people are common at crime scenes and are much harder to interpret, since the analyst must decide how many contributors are present and which peaks belong to whom, and studies circulating the same mixture to multiple laboratories have found substantial disagreement. Low template DNA, amplified from a few cells, produces stochastic effects in which peaks drop out or appear spuriously, so the result is less reliable exactly where it is most tempting to use. Transfer is the deeper problem: DNA moves between surfaces, so a person's DNA can be recovered from an object they never touched, and secondary and tertiary transfer has been demonstrated experimentally. That matters because the question a profile answers is whose DNA is present, and the question a court needs answered is how it got there, which no laboratory can determine. A widely reported Australian case saw a man charged partly on DNA later traced to contamination by paramedics who had attended both scenes.
Databases and family searching
National DNA databases hold millions of profiles and generate matches that would otherwise be impossible, including cold case resolutions decades after the offence. They also raise sustained questions. Whose profiles are retained, and for how long, has been litigated, with the European Court of Human Rights ruling in 2008 that the blanket indefinite retention of profiles from people never convicted breached privacy rights, which forced legislative change in the United Kingdom. Databases over-represent groups who are disproportionately arrested, which propagates into who is found by future searches. Familial searching looks for partial matches indicating a relative rather than the person, which has solved serious cases and effectively places the relatives of everyone on a database under surveillance without their involvement. Investigative genetic genealogy goes further, uploading a crime scene profile to consumer ancestry databases to find distant relatives and then building family trees, which identified a notorious Californian offender in 2018 and which means that a person's decision to test their own DNA has consequences for cousins who never consented.
The takeaway
Profiling counts repeats at about twenty locations chosen to be highly variable and to carry no known functional information, so a standard profile reveals nothing about health or appearance. Amplification allows work from tiny samples, which is also where reliability falls, since low template material and mixtures are much harder to interpret and laboratories disagree on them. DNA transfers between surfaces, so a profile shows whose DNA is present and never how it arrived. Database retention and familial searching raise unresolved privacy questions.