Robust Explainability: A Tutorial on Gradient-Based Attribution Methods for Deep Neural Networks
arXiv:2107.11400 · doi:10.1109/MSP.2022.3142719
Abstract
With the rise of deep neural networks, the challenge of explaining the predictions of these networks has become increasingly recognized. While many methods for explaining the decisions of deep neural networks exist, there is currently no consensus on how to evaluate them. On the other hand, robustness is a popular topic for deep learning research; however, it is hardly talked about in explainability until very recently. In this tutorial paper, we start by presenting gradient-based interpretability methods. These techniques use gradient signals to assign the burden of the decision on the input features. Later, we discuss how gradient-based methods can be evaluated for their robustness and the role that adversarial robustness plays in having meaningful explanations. We also discuss the limitations of gradient-based methods. Finally, we present the best practices and attributes that should be examined before choosing an explainability method. We conclude with the future directions for research in the area at the convergence of robustness and explainability.
23 pages, 4 figures
References in corpus (8)
- Striving for Simplicity: The All Convolutional Net
- SmoothGrad: removing noise by adding noise
- What Does Explainable AI Really Mean? A New Conceptualization of Perspectives
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- On the Connection Between Adversarial Robustness and Saliency Map Interpretability
- Bridging Adversarial Robustness and Gradient Interpretability
- Are Visual Explanations Useful? A Case Study in Model-in-the-Loop Prediction
- Contrastive Reasoning in Neural Networks
Cited by in corpus (13)
- Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research Directions
- Transformers in Time-series Analysis: A Tutorial
- Multimodal Data Integration for Oncology in the Era of Deep Neural Networks: A Review
- Trustworthy Graph Neural Networks: Aspects, Methods and Trends
- Explaining Full-disk Deep Learning Model for Solar Flare Prediction using Attribution Methods
- SoK: Modeling Explainability in Security Analytics for Interpretability, Trustworthiness, and Usability
- A Gradient Mapping Guided Explainable Deep Neural Network for Extracapsular Extension Identification in 3D Head and Neck Cancer Computed Tomography Images
- Explainable Deep Learning-based Solar Flare Prediction with post hoc Attention for Operational Forecasting
- EvalAttAI: A Holistic Approach to Evaluating Attribution Maps in Robust and Non-Robust Models
- FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
- Explainability in AI Based Applications: A Framework for Comparing Different Techniques
- On Diversity in Discriminative Neural Networks
- FREQuency ATTribution: benchmarking frequency-based occlusion for time series data