From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
arXiv:2206.03208 · doi:10.1038/S42256-023-00711-8
Abstract
The field of eXplainable Artificial Intelligence (XAI) aims to bring transparency to today's powerful but opaque deep learning models. While local XAI methods explain individual predictions in form of attribution maps, thereby identifying where important features occur (but not providing information about what they represent), global explanation techniques visualize what concepts a model has generally learned to encode. Both types of methods thus only provide partial insights and leave the burden of interpreting the model's reasoning to the user. In this work we introduce the Concept Relevance Propagation (CRP) approach, which combines the local and global perspectives and thus allows answering both the "where" and "what" questions for individual predictions. We demonstrate the capability of our method in various settings, showcasing that CRP leads to more human interpretable explanations and provides deep insights into the model's representation and reasoning through concept atlases, concept composition analyses, and quantitative investigations of concept subspaces and their role in fine-grained decision making.
87 pages (13 pages manuscript, 8 pages references, 66 pages appendix) 63 figures (6 in manuscript, 57 in appendix) 3 tables (in appendix)
References in corpus (4)
Cited by in corpus (18)
- Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research Directions
- Opening the Black-Box: A Systematic Review on Explainable AI in Remote Sensing
- Explaining Deep Learning for ECG Analysis: Building Blocks for Auditing and Knowledge Discovery
- How explainable AI affects human performance: A systematic review of the behavioural consequences of saliency maps
- Human-like object concept representations emerge naturally in multimodal large language models
- Disentangled Explanations of Neural Network Predictions by Finding Relevant Subspaces
- Concept-based Explainable Artificial Intelligence: A Survey
- Towards Interpretability in Audio and Visual Affective Machine Learning: A Review
- Towards Symbolic XAI -- Explanation Through Human Understandable Logical Relationships Between Features
- Human-Centered Evaluation of XAI Methods
- Study on the Helpfulness of Explainable Artificial Intelligence
- A three-Level Framework for LLM-Enhanced eXplainable AI: From technical explanations to natural language
- Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
- FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
- Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
- FaceX: Understanding Face Attribute Classifiers through Summary Model Explanations
- The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis
- Conceptualizing Uncertainty: A Concept-based Approach to Explaining Uncertainty