Who's responsible? Jointly quantifying the contribution of the learning algorithm and training data
arXiv:1910.04214
Abstract
A learning algorithm trained on a dataset is revealed to have poor performance on some subpopulation at test time. Where should the responsibility for this lay? It can be argued that the data is responsible, if for example training on a more representative dataset would have improved the performance. But it can similarly be argued that itself is at fault, if training a different variant on the same dataset would have improved performance. As ML becomes widespread and such failure cases more common, these types of questions are proving to be far from hypothetical. With this motivation in mind, in this work we provide a rigorous formulation of the joint credit assignment problem between a learning algorithm and a dataset . We propose Extended Shapley as a principled framework for this problem, and experiment empirically with how it can be used to address questions of ML accountability.
To appear in AAAI/ACM Conference on AI, Ethics, and Society (2021)
References in corpus (4)
Cited by in corpus (6)
- A Distributional Framework for Data Valuation
- Randomness In Neural Network Training: Characterizing The Impact of Tooling
- Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog Evaluation
- Representation Matters: Assessing the Importance of Subgroup Allocations in Training Data
- A Quantitative Perspective on Values of Domain Knowledge for Machine Learning
- Seven challenges for harmonizing explainability requirements