How a DNA Match Is Expressed and What the Number Means
A typing result reaches a report as a comparison conclusion and a figure. The conclusion says a reference is included, excluded or neither. The figure says how unusual the profile is, or how far the data favor one proposition over another, and the two are not interchangeable.

The rule in short
A short tandem repeat profile is read from an electropherogram, compared against a reference, and reported with a statistic. A random match probability estimates how often the profile would appear among unrelated people, by multiplying allele frequencies across loci under independence assumptions and a subpopulation correction. A likelihood ratio instead compares two stated propositions. Y-chromosome and mitochondrial results are lineage markers estimated by counting.
A DNA result arrives as two statements. One is a comparison conclusion: the reference profile is included, is excluded, or the comparison is inconclusive. The other is a figure, and the figure is what most of the argument turns on. The two statements come from different steps of the analysis, rest on different assumptions, and can fail independently of one another.
What the typing measures
Short tandem repeats are stretches of DNA in which a brief unit of bases repeats, and the number of repeats at a given location varies between people. An analyst extracts DNA from the sample, estimates how much is present, copies the target regions by polymerase chain reaction, and separates the copies by size in a capillary instrument. The instrument records fluorescence as the fragments pass a detector.
The record it produces is an electropherogram: a trace of peaks, each assigned a locus and a repeat number. Two peaks at a locus indicate different inherited variants. One peak indicates the same variant from both parents, or a variant whose partner failed to copy. Peak height, expressed in relative fluorescence units, carries the information used to judge whether a peak is real, whether material has been lost, and how many people contributed.
The national index rests on a designated core set of loci, and a contributing laboratory operates under published quality assurance standards, accreditation and external audits required by 34 U.S.C. § 12592. More loci means more power to discriminate between people. Whether a result appears at every locus depends on how much DNA survived, and degraded material commonly yields the smaller fragments while the larger ones fall away.
Inclusion, exclusion and the space between
An exclusion is the strongest conclusion the method offers. If the evidence profile carries a variant the reference does not have, and the analyst can rule out an artifact or an additional contributor, the reference did not contribute that material. An exclusion needs no statistic to mean something.
Inclusion is different. The phrase "cannot be excluded" says only that every variant in the reference is accounted for at the loci interpreted. It does not say the reference is the source, and standing alone it does not say how unusual the agreement is. A profile built on few loci can include very large numbers of people, and the phrase reads the same whether the agreement is remarkable or ordinary. The statistic exists to fill that gap.
Inconclusive is a third outcome, reported when the data will not support a comparison or a meaningful figure. Laboratories differ in how readily they use it, and the criteria sit in the standard operating procedure rather than in the report. An inconclusive result on a clean sample is a different event from one on a degraded trace.
Where the frequency comes from
A frequency estimate is built from allele frequency tables compiled from population samples. For each locus the analyst converts the frequencies of the observed variants into a genotype frequency, then multiplies across the loci. Multiplication assumes that the two variants at a locus are independent of each other, and that the loci are independent of one another. Both assumptions are approximations, and both hold better in a large mixed population than in a small, isolated one.
Population structure is handled by a correction that inflates the estimate to allow for shared ancestry within subgroups. The effect is conservative by design: it makes the profile look more common than raw multiplication would suggest. A laboratory also chooses which population tables to report, which is why a report often carries separate figures for several broad groups rather than one national figure.
Because those tables come from samples and not from a census, the result is an estimate carrying its own uncertainty. The National Academies report on strengthening forensic science treated nuclear DNA typing as the discipline with the firmest quantitative footing, while pressing every discipline to state the limits of what it reports. The figure is defensible, and it is still an estimate.
A random match probability is the chance that a person drawn at random from a population would have this profile. It is not the chance that the reference person is not the source, and it is not the chance that the reference person is innocent. Reversing those conditional statements is a reasoning error with a name, the prosecutor's fallacy, and it persists because the correct and the reversed version sound almost identical.
Two forms the statistic takes
The random match probability answers a frequency question: how often the profile would be expected among unrelated people. It suits a single-source profile, where there is one genotype to describe. It says nothing about relatives, who share variants at a rate set by inheritance rather than by population frequency, which is why a report sometimes carries a separate figure for siblings or other close relations.
A likelihood ratio answers a comparative question. It states how much more probable the observed data are if one proposition holds than if a stated alternative holds. That is the form used for mixed samples and for results produced by probabilistic genotyping software, because such data cannot be reduced to a single genotype. Its value depends on the two propositions chosen, so the propositions belong in the report beside the ratio.
Neither figure speaks to how the material reached the item, when it was left, or what kind of cells it came from. Those are separate questions, and the distance between a statistic about a profile and a claim about conduct is the subject of secondary transfer and the limits of touch DNA. Where more than one person contributed, the figure also rests on assumptions examined in the interpretation of mixed profiles.
| Figure reported | What it compares | How it is derived | What it cannot address |
|---|---|---|---|
| Random match probability | One profile against a population of unrelated people | Allele frequencies multiplied across loci, with a subpopulation correction | Relatives, mixed samples, and how the material arrived |
| Likelihood ratio | Two stated propositions about who contributed | Probability of the observed data under each proposition | Any proposition the analyst did not state |
| Relative frequency | The profile against a stated class of relatives | Frequencies adjusted for shared inheritance within that class | People outside the relationship the figure assumes |
| Y-chromosome haplotype frequency | The haplotype against a database of male lineages | Counting occurrences, reported with an upper confidence bound | Any distinction among men of one paternal line |
| Mitochondrial haplotype frequency | The sequence against a database of maternal lineages | Counting occurrences, reported with an upper confidence bound | Any distinction among people of one maternal line |
| Inclusion stated with no figure | Nothing quantitative | Comparison at the loci the analyst interpreted | How common the agreement actually is |
Lineage markers and counting estimates
Two marker systems break the pattern. Y-chromosome typing targets the male portion of a sample and is used where a small male contribution sits beneath a large female one. Mitochondrial typing works on material in which nuclear DNA has degraded past use, such as hair shafts without roots and old skeletal remains, because each cell carries many copies of the mitochondrial genome.
Both are lineage markers. A Y haplotype passes intact down the male line, so a father, his sons, his brothers and his paternal cousins are expected to share it. A mitochondrial sequence descends the maternal line the same way. The markers within each system travel together, so the product rule cannot be applied across them; multiplying would assume an independence that inheritance does not supply.
The frequency is therefore estimated by counting. The analyst searches a haplotype database, records how often the observed type appears, and reports an upper confidence bound that accounts for the size of that database. Such figures are far less discriminating than those from nuclear typing, and no counting estimate distinguishes members of one lineage from each other. A report that omits the limitation invites a reader to hear more than the method delivered. Whether the laboratory held accreditation, and what an external audit examined, is taken up in crime laboratory accreditation and audits.
Points to carry away
- An exclusion carries meaning without a statistic; an inclusion does not, because a partial profile can include very large numbers of people.
- A frequency estimate multiplies allele frequencies across loci and assumes independence within and between those loci.
- A subpopulation correction inflates the estimate to allow for shared ancestry, which makes the reported profile look more common than raw multiplication suggests.
- A random match probability describes a profile in a population; a likelihood ratio compares two propositions the analyst has stated.
- Transposing the statistic into a statement about the reference person is a known reasoning error rather than an interpretation of the data.
- Y-chromosome and mitochondrial types are shared across a lineage and are estimated by counting occurrences in a database with an upper confidence bound.
Questions readers ask
Why do reports list several population figures instead of one?
Allele frequencies are compiled from population samples grouped into broad categories, and the same profile can be more or less common depending on which table is used. A laboratory that reports figures for several groups is showing the range the underlying data support rather than picking the most favorable one. The practice also makes the estimate's dependence on sampling visible. None of the figures is a census, and none of them describes the particular community the item came from, which is a limitation the report should state rather than resolve.
What happens to the statistic when only some loci give results?
Every locus that fails removes a term from the multiplication, so a partial profile always produces a larger frequency than a complete one from the same person. Degraded material tends to lose the larger fragments first, which means the loci that drop out are often the same ones across cases. A laboratory reports the statistic using only the loci it interpreted, so the number and the list of loci belong together. Reading the figure without that list overstates what the comparison covered.
Can a laboratory say a profile came from one person and no one else?
Not as a matter of the underlying method. The reported figure is an estimate drawn from population samples, not a demonstration that no second person shares the profile, and close relatives share variants at a rate ordinary population tables do not capture. Some laboratories use source attribution language when the estimate passes an internal threshold, and that practice is a policy choice recorded in the standard operating procedure rather than a result the typing itself produces. The distinction is worth pressing whenever a report uses the word identification.
Sources
- National Institute of Justice — DNA Evidence: Basics of AnalyzingSets out the steps of DNA analysis, the designated core STR loci for the national index, and the roles of Y-chromosome and mitochondrial testing.
- 34 U.S.C. § 12592 — DNA identification indexConditions participation in the national index on published quality assurance standards, accreditation, and external audits not less than once every two years.
- NIST — Organization of Scientific Area Committees for Forensic ScienceMaintains the registry of discipline-specific standards, including interpretation and reporting standards a laboratory may adopt.
- National Academies — Strengthening Forensic Science in the United States: A Path ForwardAssesses the scientific footing of the forensic disciplines and calls for enforceable standards and honest statements of uncertainty.
- Federal Rule of Criminal Procedure 16 — Discovery and InspectionCovers results or reports of scientific tests and requires an expert disclosure stating the opinions, the bases and the reasons for them.
- Administrative Office of the U.S. Courts — Federal Rules of EvidenceThe adopted text of the evidence rules under which a forensic result is offered.
Premier Defense Law is a publication, not a law firm. This article states general rules and cites its sources; it is not advice about any particular case, and the law differs by state and changes over time.
More in Forensic Evidence
Fingerprint Comparison and the ACE-V Method
A latent print comparison runs through analysis, comparison, evaluation and verification. Sufficiency at the analysis stage is the examiner's judgment and is not fixed by any national minimum point count. An automated search returns ranked candidates rather than conclusions. Verification may or may not be blind. Black box testing measures the accuracy of conclusions without examining how they were reached, and a reported error rate describes study participants, not a single comparison.
Retaining a Defense Expert and Paying for One
Section 3006A(e) authorizes investigative, expert and other services necessary for adequate representation where the person is financially unable to obtain them. The application may be made ex parte and heard ex parte, so the request does not disclose the theory of the defense. Compensation is capped at an amount the court may exceed on certification approved by the chief judge of the circuit. Rule 706 supplies a court-appointed route, and Rule 16 governs disclosure once the expert testifies.
Firearms and Toolmark Identification
Firearms examination compares class, subclass and individual characteristics on fired components under a comparison microscope. The identification threshold is agreement judged sufficient by the examiner, not a fixed count of matching striae. Subclass carryover from consecutively produced tooling can imitate individual agreement. Correlation databases return ranked candidates, and conclusion wording has moved away from claims of identification to the exclusion of every other firearm.


