P-TAME: Explain Any Image Classifier with Trained Perturbations
arXiv:2501.17813 · doi:10.1109/OJSP.2025.3568756
Abstract
The adoption of Deep Neural Networks (DNNs) in critical fields where predictions need to be accompanied by justifications is hindered by their inherent black-box nature. In this paper, we introduce P-TAME (Perturbation-based Trainable Attention Mechanism for Explanations), a model-agnostic method for explaining DNN-based image classifiers. P-TAME employs an auxiliary image classifier to extract features from the input image, bypassing the need to tailor the explanation method to the internal architecture of the backbone classifier being explained. Unlike traditional perturbation-based methods, which have high computational requirements, P-TAME offers an efficient alternative by generating high-resolution explanations in a single forward pass during inference. We apply P-TAME to explain the decisions of VGG-16, ResNet-50, and ViT-B-16, three distinct and widely used image classifiers. Quantitative and qualitative results show that our method matches or outperforms previous explainability methods, including model-specific approaches. Code and trained models will be released upon acceptance.
Published in IEEE Open Journal of Signal Processing (Volume 6)
References in corpus (8)
- Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
- Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
- A Comprehensive Taxonomy for Explainable Artificial Intelligence: A Systematic Survey of Surveys on Methods and Concepts
- Making AI Intelligible: Philosophical Foundations
- Unlocking the Black Box: Analysing the EU Artificial Intelligence Act's Framework for Explainability in AI
- TAME: Attention Mechanism Based Feature Fusion for Generating Explanation Maps of Convolutional Neural Networks
- T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
- CNNs Avoid Curse of Dimensionality by Learning on Patches