When and How to Fool Explainable Models (and Humans) with Adversarial Examples
arXiv:2107.01943 · doi:10.1002/widm.1567
Abstract
Reliable deployment of machine learning models such as neural networks continues to be challenging due to several limitations. Some of the main shortcomings are the lack of interpretability and the lack of robustness against adversarial examples or out-of-distribution inputs. In this exploratory review, we explore the possibilities and limits of adversarial attacks for explainable machine learning models. First, we extend the notion of adversarial examples to fit in explainable machine learning scenarios, in which the inputs, the output classifications and the explanations of the model's decisions are assessed by humans. Next, we propose a comprehensive framework to study whether (and how) adversarial examples can be generated for explainable models under human assessment, introducing and illustrating novel attack paradigms. In particular, our framework considers a wide range of relevant yet often ignored factors such as the type of problem, the user expertise or the objective of the explanations, in order to identify the attack strategies that should be adopted in each scenario to successfully deceive the model (and the human). The intention of these contributions is to serve as a basis for a more rigorous and realistic study of adversarial examples in the field of explainable machine learning.
Updated version. 43 pages, 9 figures, 4 tables
References in corpus (23)
- Explaining and Harnessing Adversarial Examples
- Striving for Simplicity: The All Convolutional Net
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Understanding Neural Networks Through Deep Visualization
- Delving into Transferable Adversarial Examples and Black-box Attacks
- From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- Interpreting Adversarially Trained Convolutional Neural Networks
- Adversarial Images for Variational Autoencoders
- From ImageNet to Image Classification: Contextualizing Progress on Benchmarks
- Adversarial Examples for Semantic Image Segmentation
- Imperceptible Adversarial Attacks on Tabular Data
- Fairwashing Explanations with Off-Manifold Detergent
- On the Connection Between Adversarial Robustness and Saliency Map Interpretability
- Black-box adversarial attacks using Evolution Strategies
- Adversarial Attacks for Tabular Data: Application to Fraud Detection and Imbalanced Data
- Probabilistically Robust Recourse: Navigating the Trade-offs between Costs and Robustness in Algorithmic Recourse
- Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense
- A simple defense against adversarial attacks on heatmap explanations
- Foiling Explanations in Deep Neural Networks
- Counterfactual Explanations Can Be Manipulated
- Simple Transparent Adversarial Examples
- On the Veracity of Local, Model-agnostic Explanations in Audio Classification: Targeted Investigations with Adversarial Examples