It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks
arXiv:2106.05498
Abstract
Risk assessment instrument (RAI) datasets, particularly ProPublica's COMPAS dataset, are commonly used in algorithmic fairness papers due to benchmarking practices of comparing algorithms on datasets used in prior work. In many cases, this data is used as a benchmark to demonstrate good performance without accounting for the complexities of criminal justice (CJ) processes. However, we show that pretrial RAI datasets can contain numerous measurement biases and errors, and due to disparities in discretion and deployment, algorithmic fairness applied to RAI datasets is limited in making claims about real-world outcomes. These reasons make the datasets a poor fit for benchmarking under assumptions of ground truth and real-world impact. Furthermore, conventional practices of simply replicating previous data experiments may implicitly inherit or edify normative positions without explicitly interrogating value-laden assumptions. Without context of how interdisciplinary fields have engaged in CJ research and context of how RAIs operate upstream and downstream, algorithmic fairness practices are misaligned for meaningful contribution in the context of CJ, and would benefit from transparent engagement with normative considerations and values related to fairness, justice, and equality. These factors prompt questions about whether benchmarks for intrinsically socio-technical systems like the CJ system can exist in a beneficial and ethical way.
NeurIPS 2021 Datasets and Benchmarks
References in corpus (4)
Cited by in corpus (8)
- Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations
- Whither Bias Goes, I Will Go: An Integrative, Systematic Review of Algorithmic Bias Mitigation
- Can Fairness be Automated? Guidelines and Opportunities for Fairness-aware AutoML
- ESR: Ethics and Society Review of Artificial Intelligence Research
- One Model Many Scores: Using Multiverse Analysis to Prevent Fairness Hacking and Evaluate the Influence of Model Design Decisions
- Lazy Data Practices Harm Fairness Research
- Flipping the Script on Criminal Justice Risk Assessment: An actuarial model for assessing the risk the federal sentencing system poses to defendants
- Retiring Adult: New Datasets for Fair Machine Learning