A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning
arXiv:1801.05075
Abstract
Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve diverse goals in designing interpretable machine learning systems. In this paper, we propose a human attention benchmark for image and text domains using multi-layer human attention masks aggregated from multiple human annotators. We then present an evaluation study to evaluate model saliency explanations obtained using Grad-cam and LIME techniques. We demonstrate our benchmark's utility for quantitative evaluation of model explanations by comparing it with human subjective ratings and ground-truth single-layer segmentation masks evaluations. Our study results show that our threshold agnostic evaluation method with the human attention baseline is more effective than single-layer object segmentation masks to ground truth. Our experiments also reveal user biases in the subjective rating of model saliency explanations.
Benchmark Available online at https://github.com/SinaMohseni/ML-Interpretability-Evaluation-Benchmark
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Towards A Rigorous Science of Interpretable Machine Learning
- Sanity Checks for Saliency Maps
- A Benchmark for Interpretability Methods in Deep Neural Networks
- Beyond Sparsity: Tree Regularization of Deep Models for Interpretability
- Quantifying Interpretability and Trust in Machine Learning Systems
- How model accuracy and explanation fidelity influence user trust
Cited by in corpus (13)
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- Explainable Artificial Intelligence: A Survey of Needs, Techniques, Applications, and Future Direction
- Predicting Model Failure using Saliency Maps in Autonomous Driving Systems
- Pitfalls of Explainable ML: An Industry Perspective
- Evaluating neural network explanation methods using hybrid documents and morphological agreement
- A simple defense against adversarial attacks on heatmap explanations
- Aggregating explanation methods for stable and robust explainability
- Evaluating Explanation Methods for Neural Machine Translation
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- On Two XAI Cultures: A Case Study of Non-technical Explanations in Deployed AI System
- Human-grounded Evaluations of Explanation Methods for Text Classification
- Improving Attribution Methods by Learning Submodular Functions