A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness Evaluations
arXiv:2401.17486 · doi:10.1145/3613904.3642398
Abstract
Responsible design of AI systems is a shared goal across HCI and AI communities. Responsible AI (RAI) tools have been developed to support practitioners to identify, assess, and mitigate ethical issues during AI development. These tools take many forms (e.g., design playbooks, software toolkits, documentation protocols). However, research suggests that use of RAI tools is shaped by organizational contexts, raising questions about how effective such tools are in practice. To better understand how RAI tools are -- and might be -- evaluated, we conducted a qualitative analysis of 37 publications that discuss evaluations of RAI tools. We find that most evaluations focus on usability, while questions of tools' effectiveness in changing AI development are sidelined. While usability evaluations are an important approach to evaluate RAI tools, we draw on evaluation approaches from other fields to highlight developer- and community-level steps to support evaluations of RAI tools' effectiveness in shaping AI development practices and outcomes.
Accepted for publication in Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI '24), May 11--16, 2024, Honolulu, HI, USA
References in corpus (13)
- Improving fairness in machine learning systems: What do industry practitioners need?
- Evaluation and Measurement of Software Process Improvement -- A Systematic Literature Review
- BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
- Interactive Model Cards: A Human-Centered Approach to Model Documentation
- A Systematic Review and Thematic Analysis of Community-Collaborative Approaches to Computing Research
- Charting the Sociotechnical Gap in Explainable AI: A Framework to Address the Gap in XAI
- Sensible AI: Re-imagining Interpretability and Explainability using Sensemaking Theory
- WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT?
- `It is currently hodgepodge'': Examining AI/ML Practitioners' Challenges during Co-production of Responsible AI Values
- LiFT: A Scalable Framework for Measuring Fairness in ML Applications
- Amazon SageMaker Clarify: Machine Learning Bias Detection and Explainability in the Cloud
- Toward Operationalizing Pipeline-aware ML Fairness: A Research Agenda for Developing Practical Guidelines and Tools
- Exploring AI Futures Through Role Play