Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
arXiv:2503.08636 · doi:10.1007/s10994-025-06896-w
Abstract
A common belief is that intrinsically interpretable deep learning models ensure a correct, intuitive understanding of their behavior and offer greater robustness against accidental errors or intentional manipulation. However, these beliefs have not been comprehensively verified, and growing evidence casts doubt on them. In this paper, we highlight the risks related to overreliance and susceptibility to adversarial manipulation of these so-called "intrinsically (aka inherently) interpretable" models by design. We introduce two strategies for adversarial analysis with prototype manipulation and backdoor attacks against prototype-based networks, and discuss how concept bottleneck models defend against these attacks. Fooling the model's reasoning by exploiting its use of latent prototypes manifests the inherent uninterpretability of deep neural networks, leading to a false sense of security reinforced by a visual confirmation bias. The reported limitations of part-prototype networks put their trustworthiness and applicability into question, motivating further work on the robustness and alignment of (deep) interpretable models.
Accepted by Machine Learning
References in corpus (24)
- Adversarial attacks and defenses in explainable artificial intelligence: A survey
- Acquisition of Chess Knowledge in AlphaZero
- Concept Embedding Models: Beyond the Accuracy-Explainability Trade-Off
- The (de)biasing effect of GAN-based augmentation methods on skin lesion images
- This Looks Like That... Does it? Shortcomings of Latent Space Prototype Interpretability in Deep Networks
- Interpretable machine learning for time-to-event prediction in medicine and healthcare
- ProtoPFormer: Concentrating on Prototypical Parts in Vision Transformers for Interpretable Image Recognition
- Performance is not enough: the story told by a Rashomon quartet
- Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
- On the Safety of Interpretable Machine Learning: A Maximum Deviation Approach
- SAFARI: Versatile and Efficient Evaluations for Robustness of Interpretability
- This Reads Like That: Deep Learning for Interpretable Natural Language Processing
- Interpreting and Correcting Medical Image Classification with PIP-Net
- On the Robustness of Global Feature Effect Explanations
- Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype Networks
- PIPNet3D: Interpretable Detection of Alzheimer in MRI Scans
- ProtoS-ViT: Visual foundation models for sparse self-explainable classifications
- Patch-based Intuitive Multimodal Prototypes Network (PIMPNet) for Alzheimer's Disease classification
- Exploring and Interacting with the Set of Good Sparse Generalized Additive Models
- Aggregated Attributions for Explanatory Analysis of 3D Segmentation Models
- Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution
- Benchmarking the Attribution Quality of Vision Models
- When are Post-hoc Conceptual Explanations Identifiable?
- B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable