The Intriguing Relation Between Counterfactual Explanations and Adversarial Examples
arXiv:2009.05487 · doi:10.1007/s11023-021-09580-9
Abstract
The same method that creates adversarial examples (AEs) to fool image-classifiers can be used to generate counterfactual explanations (CEs) that explain algorithmic decisions. This observation has led researchers to consider CEs as AEs by another name. We argue that the relationship to the true label and the tolerance with respect to proximity are two properties that formally distinguish CEs and AEs. Based on these arguments, we introduce CEs, AEs, and related concepts mathematically in a common framework. Furthermore, we show connections between current methods for generating CEs and AEs, and estimate that the fields will merge more and more as the number of common use-cases grows.
References in corpus (20)
- Explaining and Harnessing Adversarial Examples
- Language Models are Few-Shot Learners
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- FACE: Feasible and Actionable Counterfactual Explanations
- The Pragmatic Turn in Explainable Artificial Intelligence (XAI)
- The Hidden Assumptions Behind Counterfactual Explanations and Principal Reasons
- Multi-Objective Counterfactual Explanations
- Learning Model-Agnostic Counterfactual Explanations for Tabular Data
- A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
- Interpretable Predictions of Tree-based Ensembles via Actionable Feature Tweaking
- Causality for Machine Learning
- Counterfactual State Explanations for Reinforcement Learning Agents via Generative Deep Learning
- Imperceptible Adversarial Attacks on Tabular Data
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks
- PermuteAttack: Counterfactual Explanation of Machine Learning Credit Scorecards
- Explanations of Black-Box Model Predictions by Contextual Importance and Utility
- Semantics and explanation: why counterfactual explanations produce adversarial examples in deep neural networks
- Adversarial Attacks for Tabular Data: Application to Fraud Detection and Imbalanced Data
- Natural Language Interaction with Explainable AI Models