From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI
arXiv:2201.08164 · doi:10.1145/3583558
Abstract
The rising popularity of explainable artificial intelligence (XAI) to understand high-performing black boxes raised the question of how to evaluate explanations of machine learning (ML) models. While interpretability and explainability are often presented as a subjectively validated binary property, we consider it a multi-faceted concept. We identify 12 conceptual properties, such as Compactness and Correctness, that should be evaluated for comprehensively assessing the quality of an explanation. Our so-called Co-12 properties serve as categorization scheme for systematically reviewing the evaluation practices of more than 300 papers published in the last 7 years at major AI and ML conferences that introduce an XAI method. We find that 1 in 3 papers evaluate exclusively with anecdotal evidence, and 1 in 5 papers evaluate with users. This survey also contributes to the call for objective, quantifiable evaluation methods by presenting an extensive overview of quantitative XAI evaluation methods. Our systematic collection of evaluation methods provides researchers and practitioners with concrete tools to thoroughly validate, benchmark and compare new and existing XAI methods. The Co-12 categorization scheme and our identified evaluation methods open up opportunities to include quantitative metrics as optimization criteria during model training in order to optimize for accuracy and interpretability simultaneously.
Published in ACM Computing Surveys (DOI http://dx.doi.org/10.1145/3583558). This ArXiv version includes the supplementary material. Website with categorization of XAI methods at https://utwente-dmb.github.io/xai-papers/
References in corpus (10)
- Methods for Interpreting and Understanding Deep Neural Networks
- A Survey on the Explainability of Supervised Machine Learning
- A systematic review and taxonomy of explanations in decision support and recommender systems
- XGNN: Towards Model-Level Explanations of Graph Neural Networks
- Explainability Fact Sheets: A Framework for Systematic Assessment of Explainable Approaches
- Interpretable Predictions of Tree-based Ensembles via Actionable Feature Tweaking
- A study on the Interpretability of Neural Retrieval Models using DeepSHAP
- Explainable Recommendation via Interpretable Feature Mapping and Evaluation of Explainability
- Interpretable Deep Graph Generation with Node-Edge Co-Disentanglement
- Adversarial Infidelity Learning for Model Interpretation
Cited by in corpus (49)
- Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research Directions
- Towards Human-centered Explainable AI: A Survey of User Studies for Model Explanations
- Adversarial attacks and defenses in explainable artificial intelligence: A survey
- A Comprehensive Survey on Trustworthy Graph Neural Networks: Privacy, Robustness, Fairness, and Explainability
- Opening the Black-Box: A Systematic Review on Explainable AI in Remote Sensing
- Is Conversational XAI All You Need? Human-AI Decision Making With a Conversational XAI Assistant
- A Design Framework for operationalizing Trustworthy Artificial Intelligence in Healthcare: Requirements, Tradeoffs and Challenges for its Clinical Adoption
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- Explaining Human Activity Recognition with SHAP: Validating Insights with Perturbation and Quantitative Measures
- What Does Evaluation of Explainable Artificial Intelligence Actually Tell Us? A Case for Compositional and Contextual Validation of XAI Building Blocks
- Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
- From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
- Exploring LLM-Driven Explanations for Quantum Algorithms
- Towards Directive Explanations: Crafting Explainable AI Systems for Actionable Human-AI Interactions
- Towards consistency of rule-based explainer and black box model -- fusion of rule induction and XAI-based feature importance
- Discovering robust biomarkers of psychiatric disorders from resting-state functional MRI via graph neural networks: A systematic review
- T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
- DRExplainer: Quantifiable Interpretability in Drug Response Prediction with Directed Graph Convolutional Network
- Saliency-Bench: A Comprehensive Benchmark for Evaluating Visual Explanations
- Classification Metrics for Image Explanations: Towards Building Reliable XAI-Evaluations
- Performance Measurements in the AI-Centric Computing Continuum Systems
- EvalAttAI: A Holistic Approach to Evaluating Attribution Maps in Robust and Non-Robust Models
- Interpreting Vision and Language Generative Models with Semantic Visual Priors
- Measuring Online Hate on 4chan using Pre-trained Deep Learning Models
- Evaluating the Explainability of Attributes and Prototypes for a Medical Classification Model
- Evaluating Model Explanations without Ground Truth
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- XAI-Units: Benchmarking Explainability Methods with Unit Tests
- When Can You Trust Your Explanations? A Robustness Analysis on Feature Importances
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models
- Dataset resulting from the user study on comprehensibility of explainable AI algorithms
- Beyond the Veil of Similarity: Quantifying Semantic Continuity in Explainable AI
- T-Explainer: A Model-Agnostic Explainability Framework Based on Gradients
- Self-Explanation in Social AI Agents
- The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis
- Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation
- Measuring the Effect of Background on Classification and Feature Importance in Deep Learning for AV Perception
- Growable and Interpretable Neural Control with Online Continual Learning for Autonomous Lifelong Locomotion Learning Machines
- Class-Dependent Perturbation Effects in Evaluating Time Series Attributions
- Explanation format does not matter; but explanations do -- An Eggsbert study on explaining Bayesian Optimisation tasks
- Comprehensive Evaluation of Prototype Neural Networks
- Can Offline Metrics Measure Explanation Goals? A Comparative Survey Analysis of Offline Explanation Metrics in Recommender Systems
- Hidden Conflicts in Neural Networks and Their Implications for Explainability
- From Confusion to Clarity: ProtoScore -- A Framework for Evaluating Prototype-Based XAI
- Streamlining models with explanations in the learning loop
- L'explicabilité au service de l'extraction de connaissances : application à des données médicales
- Trustworthy AI-based crack-tip segmentation using domain-guided explanations