Probabilistic Genotyping and the Software Behind It
Software now supplies the figure in most contested mixture cases. It weighs the genotype combinations that could explain a trace and reports a ratio between two propositions the analyst wrote. Both the propositions and the run settings are human inputs, and both change the answer.

The rule in short
Probabilistic genotyping weighs the genotype combinations consistent with an observed DNA profile and reports a likelihood ratio comparing two propositions. Semi-continuous models use the presence or absence of variants; fully continuous models also use peak heights. Developmental validation is performed by the developer, internal validation by the laboratory that runs it. Access disputes concern source code and the case run record, and the two are separate requests.
Probabilistic genotyping is a class of software that takes a DNA profile no analyst could resolve by hand and returns a number. It does not decide who contributed. It calculates how well the observed data are explained by one stated proposition compared with another, and reports the comparison as a ratio. Everything contestable sits either in the model or in what the analyst told the model to assume.
What the program computes
The starting point is the trace and a set of candidate explanations. For a mixture with an assumed number of contributors, there are many combinations of genotypes that could have produced the observed peaks, and the program enumerates them and assigns each a weight. The weight reflects how well that combination accounts for what the instrument recorded, given the variation expected at the amounts present.
Two propositions are then evaluated against those weights. The output is a ratio: how much more probable the observed data are under the first proposition than under the second. A large ratio favors the first, a very small one favors the second, and a ratio near unity means the data separate the propositions hardly at all. Many laboratories report the size of the ratio in verbal bands rather than reading digits aloud.
The ratio inherits everything upstream of it. The interpretation problems that make a mixture hard in the first place, set out in mixture interpretation and its limits, do not disappear because a program is doing the arithmetic; they become model assumptions instead of manual judgments.
How the two model families differ
Semi-continuous models use the presence or absence of variants. They treat a locus as a list of variants seen, estimate a probability that a variant dropped out or dropped in, and weigh genotype combinations accordingly. They make fewer assumptions about the instrument and are less sensitive to how the laboratory set its thresholds. They also discard the information carried in how tall each peak is.
Fully continuous models use the heights. They model expected peak height as a function of the amount of template contributed, model stutter as a proportion of the parent peak, and model the variability around both. Because they use more of the data, they extract more information from a difficult sample and typically produce ratios further from neutrality. The price is a larger set of parameters estimated from the laboratory's own data, and a stronger dependence on the instrument and chemistry those parameters were fitted to.
Neither family is a substitute for adequate data. Where a component is represented by a handful of low peaks, both approaches return a modest figure, and the fully continuous model returns it with more machinery. A laboratory that runs both applies written criteria to decide which model a sample goes to.
Propositions and the analyst's inputs
The ratio has no meaning apart from the two propositions compared. A common pair asks whether the sample came from the reference person and one unknown, or from two unknowns. Changing the alternative changes the answer: comparing against a relative rather than an unrelated person moves the figure, sometimes substantially. The propositions are written by a person, and should appear in the report in the words used to run the calculation.
Other inputs matter as much. The assumed number of contributors constrains the entire enumeration. An assumed contributor, where the calculation is conditioned on a known person such as the owner of a garment, removes that person's variants from the pool of things to be explained and can raise the figure for everyone else. Which references were compared, and whether the analyst knew them before fixing the interpretation, is recorded in the case file rather than in the report.
Fully continuous systems commonly sample the space of possible explanations using a Markov chain Monte Carlo method, which draws from a distribution rather than solving an equation. Repeating a run on identical data therefore returns a slightly different value. The laboratory sets a tolerance for that variation and a chain length intended to keep runs within it. A reported figure is therefore one draw from a distribution, and a rerun landing far outside the tolerance is a signal about the sample rather than a clerical difference.
The two validation studies
Developmental validation is performed by the developer of the system. It tests the model across mixture types, contributor numbers, ratios and template amounts, and establishes that the software does what its designers intended under the conditions examined. It describes the program rather than any particular laboratory that later runs it.
Internal validation is performed by the laboratory that will use the system, on its own instruments, chemistry and staff. It runs mixtures of known composition through the whole workflow and compares the figures produced against the truth known in advance. A complete study records the mixture types tested, the contributor numbers and ratios covered, the template amounts, the parameters fitted for the laboratory, the spread seen on repeated runs, and the range of conditions the laboratory concluded it could interpret. That last item is the operative one, because it is the boundary a case sample either sits inside or outside.
The NIST scientific foundation review examines probabilistic genotyping within the likelihood-ratio framework and treats the empirical record as the measure of what the tools support. Under Rule 702 the proponent must show that reliable principles were reliably applied to the facts, and the internal validation study is where the second half of that showing usually lives. The broader shape of that demonstration is taken up in the reliability showing for a forensic method.
| Record | What it contains | What it establishes |
|---|---|---|
| Developmental validation study | Developer testing across mixture types, ratios and template amounts | That the model performs as designed under the conditions tested |
| Internal validation study | The laboratory's own known-composition runs, fitted parameters and observed spread | The range of samples this laboratory demonstrated it can interpret |
| Interpretation procedure | Thresholds, contributor assignment rules, proposition wording, reporting bands | What the analyst was required to do, and what was left to judgment |
| Run settings for the case | Assumed contributors, any conditioning profile, model parameters, chain length | What the software was actually asked in this case |
| Run output and diagnostics | Per-locus values, genotype weights, convergence and replicate-run information | Whether the run behaved and how the figure was assembled |
| Source code and design documents | The implemented algorithms and the developer's specifications | How the model was built, which is separate from how it performs |
Access to the code and the record
Two requests are often conflated. The narrower one seeks case-specific material: the electronic data, the settings used, the run output, the diagnostics and the interpretation procedure. Those are laboratory records of a test performed in the case, and Rule 16 reaches results of scientific tests along with a statement of the bases for an expert's opinions. Objections tend to concern format and burden rather than entitlement.
The broader request seeks the source code. Developers resist on the ground that the code is a trade secret whose disclosure would destroy its commercial value, and they point to published validation studies as the appropriate measure of whether a tool works. The opposing position is that a figure produced by a program cannot be examined by testing outputs alone, since an implementation error is invisible in validation runs that never encounter it, and that a protective order can address commercial concerns. Courts have divided, and orders granting access commonly set conditions on who may inspect the code and under what secrecy obligations.
Practical review runs on separate tracks. A reviewer reads the internal validation study against the case sample, asks whether the settings and propositions match what the report describes, checks whether reruns fell inside the laboratory's tolerance, and asks what the figure becomes under a different contributor count. Funds for that review are available in federal court under 18 U.S.C. § 3006A(e) on an ex parte showing of necessity, described in retaining an independent forensic expert. What the figure claims, and what it does not, is the subject of how a DNA statistic is expressed.
Points to carry away
- The software reports a ratio comparing two propositions, and those propositions are written by the analyst rather than derived from the data.
- Semi-continuous models use only which variants appear; fully continuous models also model peak heights and the artifacts that accompany them.
- Developmental validation is done by the developer and internal validation by the individual laboratory, and only the second describes local conditions.
- Assumed number of contributors and any assumed contributor are inputs recorded in the run settings, not conclusions of the run.
- Repeated runs of a Markov chain Monte Carlo model return values that differ slightly, and the laboratory sets its own tolerance for that variation.
- Requests for source code and requests for the case run record are distinct, and the objections raised against each are different.
Questions readers ask
Does a high ratio mean the software worked correctly?
No. The ratio measures the relative support the data give two propositions under the model's assumptions; it says nothing about whether those assumptions fit the sample or whether the run was set up correctly. A model applied to a sample with more contributors than the laboratory validated can still return a large figure. Checking the result means checking the inputs: the assumed contributor count, the propositions, the settings used, the diagnostics from the run, and whether the sample resembled the mixtures tested in the internal validation study.
Are different programs expected to agree on the same sample?
Not exactly. Two systems modeling the same trace can return different figures because they model different features, weight genotype sets differently, and rely on different population data. On well-resolved samples the values tend to fall in the same region, and the ordinary way to compare them is by order of magnitude rather than by digits. Divergence widens on low-template and high-contributor samples, which is another way of saying the models agree where interpretation is easy and separate where it is hard.
What is the difference between a model assumption and a case assumption?
A model assumption is built into the software: how stutter is modeled, how peak heights are expected to vary, how allele frequencies are drawn. It applies to every case the program processes and is examined in validation. A case assumption is entered by the analyst for one sample, such as the number of contributors, whether a known person is treated as present, and which propositions are compared. Model assumptions are tested by studies; case assumptions are tested only by asking what would happen if they changed.
Sources
- NIST — DNA Mixture Interpretation: A NIST Scientific Foundation ReviewExamines probabilistic genotyping within a likelihood-ratio framework and the empirical support for its use on complex mixtures.
- Federal Rule of Evidence 702 — Testimony by Expert WitnessesRequires the proponent to show that the method rests on sufficient facts, reliable principles, and a reliable application of those principles to the facts.
- Federal Rule of Criminal Procedure 16 — Discovery and InspectionCovers results and reports of scientific tests and requires a written expert disclosure stating the opinions and the bases and reasons for them.
- 18 U.S.C. § 3006A — Adequate representation of defendantsSubsection (e) authorizes services other than counsel on an ex parte showing of necessity, with a capped amount the court may exceed on certification.
- 34 U.S.C. § 12592 — DNA identification indexTies participation in the national index to published quality assurance standards, accreditation, and periodic external audits.
- NIST — Organization of Scientific Area Committees for Forensic ScienceMaintains the registry of standards, including those addressing validation of interpretation systems and the content of case records.
Premier Defense Law is a publication, not a law firm. This article states general rules and cites its sources; it is not advice about any particular case, and the law differs by state and changes over time.
More in Forensic Evidence
Fingerprint Comparison and the ACE-V Method
A latent print comparison runs through analysis, comparison, evaluation and verification. Sufficiency at the analysis stage is the examiner's judgment and is not fixed by any national minimum point count. An automated search returns ranked candidates rather than conclusions. Verification may or may not be blind. Black box testing measures the accuracy of conclusions without examining how they were reached, and a reported error rate describes study participants, not a single comparison.
Retaining a Defense Expert and Paying for One
Section 3006A(e) authorizes investigative, expert and other services necessary for adequate representation where the person is financially unable to obtain them. The application may be made ex parte and heard ex parte, so the request does not disclose the theory of the defense. Compensation is capped at an amount the court may exceed on certification approved by the chief judge of the circuit. Rule 706 supplies a court-appointed route, and Rule 16 governs disclosure once the expert testifies.
Firearms and Toolmark Identification
Firearms examination compares class, subclass and individual characteristics on fired components under a comparison microscope. The identification threshold is agreement judged sufficient by the examiner, not a fixed count of matching striae. Subclass carryover from consecutively produced tooling can imitate individual agreement. Correlation databases return ranked candidates, and conclusion wording has moved away from claims of identification to the exclusion of every other firearm.


