Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
arXiv:2502.00428 · doi:10.1145/3706598.3713963
Abstract
Independent algorithm audits hold the promise of bringing accountability to automated decision-making. However, third-party audits are often hindered by access restrictions, forcing auditors to rely on limited, low-quality data. To study how these limitations impact research integrity, we conduct audit simulations on two realistic case studies for recidivism and healthcare coverage prediction. We examine the accuracy of estimating group parity metrics across three levels of access: (a) aggregated statistics, (b) individual-level data with model outputs, and (c) individual-level data without model outputs. Despite selecting one of the simplest tasks for algorithmic auditing, we find that data minimization and anonymization practices can strongly increase error rates on individual-level data, leading to unreliable assessments. We discuss implications for independent auditors, as well as potential avenues for HCI researchers and regulators to improve data access and enable both reliable and holistic evaluations.
30 pages, 12 figures. To be published in CHI Conference on Human Factors in Computing Systems (CHI '25)
References in corpus (8)
- Improving fairness in machine learning systems: What do industry practitioners need?
- 'It's Reducing a Human Being to a Percentage'; Perceptions of Justice in Algorithmic Decisions
- Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem
- Digital welfare fraud detection and the Dutch SyRI judgment
- Diffprivlib: The IBM Differential Privacy Library
- Sociotechnical Audits: Broadening the Algorithm Auditing Lens to Investigate Targeted Advertising
- A Safe Harbor for AI Evaluation and Red Teaming
- Beyond Fairness Metrics: Roadblocks and Challenges for Ethical AI in Practice