The Gun That Only One Bullet Fits
When a bullet is recovered from a body, a firearms examiner looks at it under a microscope, looks at a test bullet fired from the suspect’s gun, and traditionally tells the jury the two came from the same weapon, to a practical certainty. That testimony has helped decide murder cases for a hundred years. Over the last decade, and with real force since 2023, a coalition of statisticians and legal scholars has argued that the number underneath the certainty is unknown, and that the discipline’s celebrated sub-one-percent error rate is an artifact of how its own validation studies are scored.[10][6] In 2023 a state supreme court agreed enough to forbid the certainty.[2] The fight is worth an outsider’s attention because it is a rare, clean case of a mature expert practice being audited in public by the tools of another field, and losing ground.
The field in brief
When a gun fires, its internal surfaces stamp the ammunition. The rifling in the barrel cuts spiral grooves into the bullet, and the breech face, firing pin, and extractor scrape and dent the cartridge case. Machining leaves each surface faintly and randomly imperfect, so the marks it transfers are, the theory holds, effectively unique to that weapon. A toolmark is any such impression left by a harder object on a softer one, and firearms identification is the special case where the tool is a gun.
The examiner’s instrument is the comparison microscope, two scopes joined by an optical bridge so the questioned bullet and a test-fired reference can be viewed touching, side by side, in a single field of view. The examiner rotates them looking for individual characteristics, the fine accidental striations and impressions, as opposed to class characteristics like caliber or the number of rifling grooves, which every gun of a model shares. The governing doctrine, the AFTE Theory of Identification, named for the Association of Firearm and Tool Mark Examiners, permits a declaration of common origin when the examiner sees sufficient agreement.[4] The association defines that, in its own words, as agreement so extensive that the likelihood another tool could have made the mark is a practical impossibility, while conceding that the judgment is subjective in nature, founded on scientific principles and based on the examiner’s training and experience.[4] That concession, subjective by the field’s own hand, is where the trouble starts.
The discipline reports its conclusions on a range that runs from identification through inconclusive to elimination, and the inconclusive is used when the examiner will not commit either way. Hold that word, because almost everything in the fight turns on what an inconclusive answer means and how it is counted.
The evidentiary standard behind all of this is not the field’s alone. In American courts, scientific testimony must clear the Daubert bar, a judge acting as gatekeeper who asks whether a method has a known error rate and has been empirically tested, not merely whether its practitioners believe in it. In 2016 the President’s Council of Advisors on Science and Technology, known as PCAST, applied that lens and coined a phrase now central to the dispute, foundational validity, meaning a method has been shown by well-designed studies to work, with a measured error rate, before anyone applies it to a case.[1]
The fight
PCAST lit the fuse. Its 2016 report went through the firearms literature and found that only a single black-box study, one in which examiners judge samples of known ground truth without being told the answers, had been designed well enough to count.[1] That lone study, run by the Ames Laboratory for the Defense Department, put the false-positive rate near one in sixty-six. The verdict was not that the number was damning but that one study cannot establish a field, because foundational validity requires reproducibility, and firearms identification had not been shown to have it.[1]
The discipline answered with more studies, most prominently a 2022 follow-up known as Ames II, in which 173 examiners worked through thousands of hard comparisons and still produced false-positive rates below one percent.[11] To defenders the case was closed. In 2023 a group of psychologists led by Max Guyll and Stephanie Madon published in a leading science journal a fresh validation of cartridge-case comparison, reporting high sensitivity and specificity and arguing that examiners were, simply, accurate.[8]
Then the statisticians opened the hood. Their objection is arithmetical rather than rhetorical, and it centers on the inconclusive. In these studies, when an examiner declines to call a match that truly exists, the response is frequently scored as correct, or dropped from the denominator entirely, rather than counted as a miss. Because examiners under test return inconclusive far more often than they do in real casework, this scoring quietly drains errors out of the tally. Reclassify inconclusives as the failures-to-identify they sometimes are, the critics show, and the advertised sub-one-percent rate swells, with one group putting the recomputed figure for the original Ames data around thirty-five percent and other reworkings running higher.[6] The same critics noted that in Ames II an examiner agreed with their own earlier call only about two-thirds of the time, and two examiners agreed with each other less than a third of the time, thin ground for a practical impossibility.[6]
The exchange turned technical and personal in print. Michael Rosenblum, a Johns Hopkins biostatistician, joined by Maria Cuellar of Penn, Susan Vanderplas of Nebraska, William Thompson of UC Irvine, and others, filed a reply arguing that the defenders had baked in an equiprobability assumption, treating same gun and different gun as equally likely before any evidence, that mathematically inflates the apparent strength of a match against a defendant.[9] Alicia Carriquiry, who directs the federally funded forensic-statistics center CSAFE at Iowa State, filed her own objection alongside it.[9] In 2024 Cuellar, Vanderplas, Amanda Luby, and Rosenblum went further, publishing a study that examined all twenty-eight black-box firearms studies and concluded every one carried flaws grave enough that, once corrected, the confidence intervals on error could stretch past fifty percent.[10] Their claim is not that examiners are usually wrong. It is that the field has never measured how often they are, and cannot say.
The moment it broke into law came in June 2023. In Abruquah v. State, the Maryland Supreme Court split four to three and, with Chief Justice Matthew Fader writing, overturned a murder conviction, holding that firearms identification has not been shown to reliably link a particular bullet to a particular firearm.[2][3] An examiner in Maryland may now testify that a bullet is consistent with having been fired from the gun, but not that it was. The association’s rejoinder, published to its members, was that the court had misread the science and had mandated consistent-with language that no error-rate study has ever tested.[5] Prosecutors mounted a public defense, with Raymond Valerio of the Queens County District Attorney’s office and Nelson Bunn of the National District Attorneys Association arguing in print that Ames II and later work had met the PCAST bar and that inconclusives are honest limits rather than hidden errors.[7]
Where it stands is unsettled. Maryland’s appellate courts were still working out the reach of the ruling into 2026, other states remain split, and the newest blow came from an unexpected direction.[13] In a 2025 study of casework at the Houston Forensic Science Center, Nicholas Scurich found examiners were about 43.5 percent more likely to return inconclusive once they realized they were being tested, direct evidence that the validation studies measure behaviour under observation rather than behaviour in the room where verdicts are made.[12]
What the fight reveals
Strip away the courtroom drama and the dispute is about a single accounting choice, and on that narrow question the evidence now favors the critics. The reported error rates of firearms identification depend heavily on treating inconclusive answers as something other than errors, and once that treatment is examined, by statisticians and by the Houston finding that the inconclusive itself moves when examiners feel watched, the headline numbers stop meaning what they are used in court to mean.[10][12] That much is established. It is why a state supreme court, applying a standard imported from outside forensics, drew the line it did.[2]
What the dispute exposes about the field is structural. Firearms identification grew up as a craft inside crime labs, validated by consensus among practitioners rather than by adversarial statistical audit, and the theory’s frank admission of subjectivity was never a problem until a neighbouring discipline showed up asking for the error bars. The field is now being graded by people who did not train in it, using standards it did not set, and it is discovering that ‘we are almost never wrong’ and ‘we have measured how often we are wrong’ are different sentences.
The genuinely unresolved thing is the number itself. The critics have shown the reported error rate is unsound, yet they have not shown what the true rate is, and they say so honestly. It could be low. The scandal the statisticians have proven is not that examiners are frequently wrong. It is that, after a century in the witness box, no one can yet tell you how often they are.
- President’s Council of Advisors on Science and Technology, Forensic Science in Criminal Courts (2016) — denied firearms identification foundational validity and found only one adequately designed black-box study.
- Abruquah v. State, Maryland Supreme Court (2023) — the four-to-three ruling barring unqualified match testimony.
- Reason, on the Maryland court limiting bullet-matching testimony (2023) — a summary of the Abruquah holding; field press.
- AFTE Theory of Identification — the discipline’s own doctrine of sufficient agreement and its admission of subjectivity.
- Association of Firearm and Tool Mark Examiners, response to the Abruquah decision (2024) — the association’s official rebuttal to the court.
- Faigman, Scurich and Albright, ‘The Field of Firearms Forensics Is Flawed,’ Scientific American — the critics’ argument on inconclusives and repeatability.
- Valerio and Bunn, ‘Firearm Forensics Has Proven Reliable,’ Scientific American — the prosecutors’ defense of reliability.
- Guyll, Madon, Scherr, Lynch and colleagues, validity of forensic cartridge-case comparisons, PNAS (2023) — the defenders’ validation study.
- Rosenblum, Cuellar, Vanderplas, Thompson and colleagues, on incorrect statistical reasoning in Guyll et al., PNAS (2024), with a companion letter by Carriquiry and Ommen — the statistical rebuttal.
- Cuellar, Vanderplas, Luby and Rosenblum, methodological problems in every black-box study of forensic firearm comparisons, Law, Probability and Risk (2024).
- Monson, Smith and Peters, accuracy of comparison decisions by forensic firearms examiners (Ames II), Journal of Forensic Sciences (2022) — the field’s largest validation study.
- Naked Capitalism, ‘Rifling Through the Evidence’ (2026) — reports the 2025 Houston Forensic Science Center finding on examiner behaviour under observation and the current state of the fight.
- State v. Thornton & Dunbar, Maryland (2026) — an appellate decision applying Abruquah, showing the fight is still live.