Crime Solvers Central
CSC
291 Cases Solved. Advancing justice for missing persons, unsolved homicides, unidentified and unclaimed remains.
Back to Articles

8 Fields to Publish Cold Case Racial Data for Researchers & Advocates

Published: October 01, 2026

8 Fields to Publish Cold Case Racial Data for Researchers & Advocates

Researcher comparing public crime data sources

Start with the National Registry of Exonerations, NamUs bi-annual reports, and Department of Justice cold-case and Till Act materials: these are the primary published sources for racial disparity data in cold cases and wrongful convictions. Download the underlying datasets or request them directly, then check each one’s race and ethnicity fields before citing a figure, since coverage and reporting practices vary widely across sources.


TL;DR:

  • Racial disparity data from sources like the National Registry of Exonerations and NamUs are limited by inconsistent reporting and incomplete fields, especially for race and ethnicity.
  • Most datasets reflect only cases reported or publicly known, which introduces selection and reporting biases that can distort true disparities.
  • Proper analysis requires clear case identifiers, standardized race coding, and disclosure of denominator definitions for meaningful rate calculations.
  • Publication of race and ethnicity data should follow strict privacy and ethics protocols, including minimum cell sizes and consultation with affected communities.
  • Cross-checking findings across multiple sources and applying sensitivity tests enhances confidence in racial disparity assessments in cold cases.

Crimesolverscentral
Explore Cases Behind the Data
Crime Solvers Central catalogs missing persons and unsolved homicides nationwide, helping researchers and advocates examine cases connected to justice and closure.
Explore the case database

Table of Contents

Key repositories and reports: where published datasets and cold-case reports live

Four sources anchor most credible work on racial disparity in cold cases and exonerations. Each publishes different information, updates on a different schedule, and requires a different access method.

The National Registry of Exonerations tracks known exonerations across the country and lets researchers query cases by race, crime type, and contributing factor such as official misconduct or mistaken identification. Its analysis found that Black people accounted for a disproportionately high share of exonerations relative to their population share as of 2022, a gap the registry ties to differing causes of wrongful conviction across crime categories.

NamUs publishes bi-annual reports alongside a searchable case database covering missing, unidentified, and unclaimed persons. The reports break these categories down by race and ethnicity and by how long a case has been open, but NamUs is explicit that its numbers reflect only cases entered into the system, not a full national count.

The DOJ Civil Rights Division maintains cold-case materials tied to the Emmett Till Act, including closing memoranda that document why a case was opened, investigated, and ultimately closed. Operation Lady Justice outputs and related Indian country reporting extend this transparency effort to missing and murdered Indigenous people cases, an area where federal data has historically been thin.

Beyond these three, researchers have other channels worth checking:

  • The National Institute of Justice publishes crime and justice data portals that sometimes host cold-case related datasets.
  • FBI Indian country reporting covers tribal jurisdictions with separate reporting requirements from state and local agencies.
  • Many state and local agencies post cold-case lists or annual reports on their own websites, accessible directly or through a public records request.
  • FOIA requests remain the fallback when a dataset exists but hasn’t been published in a usable format.

Congressional report language accompanying cold-case legislation, including House Report 117-280, directs the National Institute of Justice to publish disaggregated cold-case murder statistics annually. That requirement, once fully implemented, should make year-over-year comparisons easier than they are today.

What each major dataset contains and its limits

Before building an analysis, know what variables you can actually pull from each source and where the gaps sit.

Typical fields across these datasets include:

  • Victim and offender race and ethnicity, though completeness varies by source and by year.
  • Victim or offender role in the case, alongside offense type and incident date.
  • Jurisdiction, agency, and case status or clearance outcome.
  • Contributing factors in exonerations, such as official misconduct, false confession, or mistaken identification.

None of these sources give you a clean, fully representative national count. NamUs notes explicitly that its reports reflect only cases reported into the system, and a later report adds that race and ethnicity fields reflect whatever the reporting party entered, which carries its own unknown bias. The National Registry of Exonerations depends on cases becoming publicly known, so wrongful convictions that never surface, especially in jurisdictions with weak defense resources, don’t appear in the count.

DOJ’s own cold-case materials illustrate a related problem: administrative closure. A case can be marked closed because a witness died or evidence no longer supports prosecution, not because it was solved. Treating “closed” as equivalent to “resolved” will distort any clearance-rate comparison built on these records. Indian country reporting shows a similar pattern, with a high share of cases closed for insufficient evidence rather than identified outcomes. Any dataset you pull from one of these sources needs a close read of its codebook before you treat a field as comparable to the same-named field elsewhere.

Minimal standard data fields and disaggregation checklist for publishable analyses

A publishable analysis needs a consistent field structure so other researchers can replicate or extend it. Use this as a working checklist before release:

  1. Case identifier that allows tracking without exposing personally identifying details in a public release.
  2. Incident date and, where relevant, date of case closure or clearance.
  3. Jurisdiction and agency responsible for the case.
  4. Victim race and ethnicity, coded to a standard schema rather than a free-text field.
  5. Offender race and ethnicity, when known, with a clear “unknown or unrecorded” category rather than leaving it blank.
  6. Offense type or case category, distinguishing homicide, missing person, and exoneration where applicable.
  7. Case status or clearance outcome, with administrative closure flagged separately from an identified resolution.
  8. Time to clearance or time to exoneration, calculated consistently across cases.

Race and ethnicity coding should follow a fixed set of categories with an explicit “unknown/missing” option rather than defaulting missing data to a majority category, a practice that quietly inflates apparent disparities or erases them depending on which group gets the default. For any table broken down by race, jurisdiction, or case type, suppress cells below a minimum threshold, commonly five, to avoid identifying individuals in small jurisdictions. Every release needs an accompanying codebook defining each field, its source, and its known limitations, along with metadata on when the data was pulled and from which version of the source.

Rate calculations deserve particular care. A clearance rate is only meaningful when you state the denominator, whether that’s all reported cases, all cases opened in a given year, or all cases meeting some other criteria, since switching denominators can flip a comparison entirely.

Illustration comparing different data denominators

Pro Tip: Publish your denominator definition in the same sentence as any rate you report, not in a footnote three pages later.

How to publish responsibly: privacy, ethics, and stakeholder engagement

Publishing racial disparity data on cold cases carries real risk to the families involved, and the decision between aggregated and case-level release should come early, not as an afterthought.

  • Default to aggregated tables when case-level detail could identify a living relative, reveal sensitive medical information, or expose an active lead.
  • Apply a minimum cell-size threshold, commonly five, to any breakdown by race, jurisdiction, or case type.
  • Consult affected communities and, where a case involves a tribal jurisdiction, the relevant tribal authority before publication.
  • Notify families when a release could bring renewed public attention to a specific unsolved case.
  • Run the analysis through an institutional review process or equivalent ethics check when working with identifiable records, and confirm what a FOIA response can and cannot legally include before treating it as complete.

Document where every figure originated, keep a reproducible codebook alongside the published data, and write a limitations section that states plainly what the data can’t tell you. A report that hides its blind spots erodes trust faster than one that names them upfront.

Pro Tip: Before any public release, run a harm assessment asking whether the data could re-victimize a family or expose an ongoing investigative lead. If the answer is unclear, publish the aggregate and hold the case-level detail.

Interpreting disparities: common biases and robustness checks

A racial disparity in the raw numbers isn’t automatically evidence of bias in the system that produced them, and separating the two takes deliberate checking.

Selection bias affects registry-based counts directly: the National Registry of Exonerations can only document exonerations that become known, and access to the legal resources needed to secure one is not distributed evenly. Reporting bias works similarly in NamUs data, where entries depend on what a reporting party chooses to record. Administrative closure adds another layer of distortion: a DOJ Indian country report shows a high share of cases closed for insufficient evidence rather than resolved, which can make a jurisdiction look more “efficient” than it actually is if closures are counted as clearances.

Cross-race identification problems show up repeatedly in exoneration data as a contributing factor to wrongful conviction, and official misconduct patterns differ by race and by crime type, according to the registry’s own analysis. Community-level factors matter too: a study using Indianapolis data from 2007 to 2017 found that resident complaints about neighborhood disorder were associated with lower odds of case clearance, a finding that intersects with race and socioeconomic status without being a simple product of either.

To build confidence in a disparity finding, run these checks:

  • Stratify comparisons by crime type and jurisdiction rather than pooling everything into one national number.
  • Test sensitivity to different denominator choices before publishing a single rate.
  • Compare results across at least two independent sources when the topic allows it, since a pattern that holds in both NamUs and DOJ data is harder to dismiss as an artifact of one system’s reporting quirks.

How Crime Solvers Central’s data and resources can support publication and advocacy

Public datasets rarely cover every case, and that’s where a resource like the cold case database by state can fill in detail. Crime Solvers Central maintains a searchable national database of cold cases categorized by state and case type, which researchers can use to cross-check gaps in federal reporting or generate leads for a FOIA request when a public agency hasn’t published full case detail. Its community and volunteer coordination tools can also support the outreach work that responsible publication requires, particularly when a project needs to reach families or local advocates before a release.

Treat this database as a supplement, not a substitute for the primary sources above. Verify provenance and update cadence before citing any figure from it, the same standard you’d apply to NamUs or the registry. For guidance on using a national database alongside these public sources, see how to leverage a national database of cold cases.

In the next 30 days, download the core datasets from the National Registry of Exonerations and NamUs, check how complete their race and ethnicity fields actually are, and run baseline descriptive tables before drawing conclusions. In the next 90 days, file FOIA requests for any missing fields, consult with affected families or tribal authorities if your analysis touches their cases, and draft an ethics statement to accompany any public release. Treat the checklist as sequential: description before interpretation, consultation before publication.

— Crime

Sources

FAQ

What is the National Registry of Exonerations used for?

It’s a searchable database of known exonerations in the country, letting researchers filter by race, crime type, and contributing factor such as official misconduct or mistaken identification. Its analysis found Black people accounted for about 53% of exonerations as of 2022 against 13.6% of the population, making it a primary source for racial disparity research in wrongful convictions.

How often does NamUs publish its case reports?

NamUs publishes bi-annual reports covering unresolved missing, unidentified, and unclaimed persons cases, broken down by race, ethnicity, and case age. The reports state clearly that they reflect only cases entered into the system, not a complete national count.

Why do cold-case race and ethnicity fields have so many gaps?

Fields like race and ethnicity in these databases reflect whatever the reporting party chose to enter, and NamUs itself warns that this introduces unknown bias into any national pattern drawn from the data. Researchers should treat these fields as a starting point for investigation, not a finished count, and document the gap in any published analysis.

What’s the minimum cell size for publishing disaggregated case data?

A common practice is to suppress any breakdown cell below five cases to avoid identifying individuals in small jurisdictions. This threshold should be stated explicitly in the accompanying codebook alongside the source and date of the underlying data.

Can Crime Solvers Central’s database be cited as a primary source?

Its cold case database, covering a large number of cases by state and type, can supplement public datasets like NamUs or the National Registry of Exonerations, particularly for filling regional gaps or generating FOIA leads. Researchers should still verify its provenance and update cadence before citing a specific figure, the same standard applied to any government source.