Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
arXiv:1710.11063 · doi:10.1109/WACV.2018.00097
Abstract
Over the last decade, Convolutional Neural Network (CNN) models have been highly successful in solving complex vision problems. However, these deep models are perceived as "black box" methods considering the lack of understanding of their internal functioning. There has been a significant recent interest in developing explainable deep learning models, and this paper is an effort in this direction. Building on a recently proposed method called Grad-CAM, we propose a generalized method called Grad-CAM++ that can provide better visual explanations of CNN model predictions, in terms of better object localization as well as explaining occurrences of multiple object instances in a single image, when compared to state-of-the-art. We provide a mathematical derivation for the proposed method, which uses a weighted combination of the positive partial derivatives of the last convolutional layer feature maps with respect to a specific class score as weights to generate a visual explanation for the corresponding class label. Our extensive experiments and evaluations, both subjective and objective, on standard datasets showed that Grad-CAM++ provides promising human-interpretable visual explanations for a given CNN architecture across multiple tasks including classification, image caption generation and 3D action recognition; as well as in new settings such as knowledge distillation.
17 Pages, 15 Figures, 11 Tables. Accepted in the proceedings of IEEE Winter Conf. on Applications of Computer Vision (WACV2018). Extended version is under review at IEEE Transactions on Pattern Analysis and Machine Intelligence
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Striving for Simplicity: The All Convolutional Net
- FitNets: Hints for Thin Deep Nets
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Understanding Neural Networks Through Deep Visualization
- Object Detectors Emerge in Deep Scene CNNs
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- Tell Me Where to Look: Guided Attention Inference Network
- Interpretable Learning for Self-Driving Cars by Visualizing Causal Attention
Cited by in corpus (269)
- Deep Neural Networks and Tabular Data: A Survey
- Reliable Tuberculosis Detection using Chest X-ray with Deep Learning, Segmentation and Visualization
- Automatic Classification of Defective Photovoltaic Module Cells in Electroluminescence Images
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- Eigen-CAM: Class Activation Map using Principal Components
- Explanations in Autonomous Driving: A Survey
- Graph-Based Deep Learning for Medical Diagnosis and Analysis: Past, Present and Future
- An Empirical Study of Remote Sensing Pretraining
- VLocNet++: Deep Multitask Learning for Semantic Visual Localization and Odometry
- Ablation Studies in Artificial Neural Networks
- Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models
- A Survey on Graph-Based Deep Learning for Computational Histopathology
- Learning Hierarchical Attention for Weakly-supervised Chest X-Ray Abnormality Localization and Diagnosis
- Explainable artificial intelligence in breast cancer detection and risk prediction: A systematic scoping review
- COV-ECGNET: COVID-19 detection using ECG trace images with deep convolutional neural network
- ULSAM: Ultra-Lightweight Subspace Attention Module for Compact Convolutional Neural Networks
- Review: Deep Learning in Electron Microscopy
- Computer Vision Tool for Detection, Mapping and Fault Classification of PV Modules in Aerial IR Videos
- One-Class Classification: A Survey
- Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks
- Domain Generalization for Medical Image Analysis: A Review
- Explainable Diabetic Retinopathy Detection and Retinal Image Generation
- Attention-embedded Quadratic Network (Qttention) for Effective and Interpretable Bearing Fault Diagnosis
- AI Security for Geoscience and Remote Sensing: Challenges and Future Trends
- Implementing local-explainability in Gradient Boosting Trees: Feature Contribution
- Dynamic Gesture Recognition by Using CNNs and Star RGB: a Temporal Information Condensation
- PDCOVIDNet: A Parallel-Dilated Convolutional Neural Network Architecture for Detecting COVID-19 from Chest X-Ray Images
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
- Ensemble manifold based regularized multi-modal graph convolutional network for cognitive ability prediction
- Learning to Disentangle Scenes for Person Re-identification
- ECOVNet: An Ensemble of Deep Convolutional Neural Networks Based on EfficientNet to Detect COVID-19 From Chest X-rays
- Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation
- Rotate to Attend: Convolutional Triplet Attention Module
- SS-CAM: Smoothed Score-CAM for Sharper Visual Feature Localization
- Age-Net: An MRI-Based Iterative Framework for Brain Biological Age Estimation
- Simultaneously Localize, Segment and Rank the Camouflaged Objects
- Efficient Saliency Maps for Explainable AI
- Deep Weakly-Supervised Learning Methods for Classification and Localization in Histology Images: A Survey
- COVID CT-Net: Predicting Covid-19 From Chest CT Images Using Attentional Convolutional Network
- Driving Behavior Explanation with Multi-level Fusion
- GaNDLF: A Generally Nuanced Deep Learning Framework for Scalable End-to-End Clinical Workflows in Medical Imaging
- Interpreting Adversarial Examples by Activation Promotion and Suppression
- Visual Explanation by Interpretation: Improving Visual Feedback Capabilities of Deep Neural Networks
- Fingerprint Presentation Attack Detector Using Global-Local Model
- Saliency Tubes: Visual Explanations for Spatio-Temporal Convolutions
- Attention Branch Network: Learning of Attention Mechanism for Visual Explanation
- Group-CAM: Group Score-Weighted Visual Explanations for Deep Convolutional Networks
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution
- CHS-Net: A Deep learning approach for hierarchical segmentation of COVID-19 infected CT images
- Domain-invariant Similarity Activation Map Contrastive Learning for Retrieval-based Long-term Visual Localization
- IS-CAM: Integrated Score-CAM for axiomatic-based explanations
- Interpretative Computer-aided Lung Cancer Diagnosis: from Radiology Analysis to Malignancy Evaluation
- Explainable Image Similarity: Integrating Siamese Networks and Grad-CAM
- SLAP: Improving Physical Adversarial Examples with Short-Lived Adversarial Perturbations
- ES-Net: Erasing Salient Parts to Learn More in Re-Identification
- Explaining Clinical Decision Support Systems in Medical Imaging using Cycle-Consistent Activation Maximization
- Counterfactual Explanation Based on Gradual Construction for Deep Networks
- BayLIME: Bayesian Local Interpretable Model-Agnostic Explanations
- Entanglement-guided architectures of machine learning by quantum tensor network
- Explaining deep learning of galaxy morphology with saliency mapping
- The effectiveness of feature attribution methods and its correlation with automatic evaluation scores
- Smoothed Geometry for Robust Attribution
- EARL: An Elliptical Distribution aided Adaptive Rotation Label Assignment for Oriented Object Detection in Remote Sensing Images
- A deep learning model for burn depth classification using ultrasound imaging
- Embedding Deep Networks into Visual Explanations
- M3d-CAM: A PyTorch library to generate 3D data attention maps for medical deep learning
- ThoraX-PriorNet: A Novel Attention-Based Architecture Using Anatomical Prior Probability Maps for Thoracic Disease Classification
- Towards Visually Explaining Variational Autoencoders
- Explainable Model-Agnostic Similarity and Confidence in Face Verification
- Explaining Full-disk Deep Learning Model for Solar Flare Prediction using Attribution Methods
- xCos: An Explainable Cosine Metric for Face Verification Task
- Synthetic Benchmarks for Scientific Research in Explainable Machine Learning
- XAIport: A Service Framework for the Early Adoption of XAI in AI Model Development
- Interpreting Black-box Machine Learning Models for High Dimensional Datasets
- Explaining Knowledge Distillation by Quantifying the Knowledge
- On the Black-box Explainability of Object Detection Models for Safe and Trustworthy Industrial Applications
- Weakly-Supervised Action Localization and Action Recognition using Global-Local Attention of 3D CNN
- Learning at a Glance: Towards Interpretable Data-limited Continual Semantic Segmentation via Semantic-Invariance Modelling
- IAIA-BL: A Case-based Interpretable Deep Learning Model for Classification of Mass Lesions in Digital Mammography
- ViGAT: Bottom-up event recognition and explanation in video using factorized graph attention network
- A Novel Disaster Image Dataset and Characteristics Analysis using Attention Model
- Evaluating Visual Explanations of Attention Maps for Transformer-based Medical Imaging
- Exploring the Efficacy of Base Data Augmentation Methods in Deep Learning-Based Radiograph Classification of Knee Joint Osteoarthritis
- Deeply Explain CNN via Hierarchical Decomposition
- An Explainable Contrastive-based Dilated Convolutional Network with Transformer for Pediatric Pneumonia Detection
- Occlusion Sensitivity Analysis with Augmentation Subspace Perturbation in Deep Feature Space
- Saliency Cards: A Framework to Characterize and Compare Saliency Methods
- Visual Explanation for Deep Metric Learning
- FathomNet: An underwater image training database for ocean exploration and discovery
- Black-box Error Diagnosis in Deep Neural Networks for Computer Vision: a Survey of Tools
- Dual Attention Suppression Attack: Generate Adversarial Camouflage in Physical World
- Towards Interpretability in Audio and Visual Affective Machine Learning: A Review
- An Open API Architecture to Discover the Trustworthy Explanation of Cloud AI Services
- Spectral decoupling allows training transferable neural networks in medical imaging
- Explainable Deep Learning-based Solar Flare Prediction with post hoc Attention for Operational Forecasting
- Local Black-box Adversarial Attacks: A Query Efficient Approach
- Accurate Explanation Model for Image Classifiers using Class Association Embedding
- Brain Stroke Detection and Classification Using CT Imaging with Transformer Models and Explainable AI
- Hierarchical Contrastive Motion Learning for Video Action Recognition
- Strategy to Increase the Safety of a DNN-based Perception for HAD Systems
- Optimising Knee Injury Detection with Spatial Attention and Validating Localisation Ability
- PCAMs: Weakly Supervised Semantic Segmentation Using Point Supervision
- Visualizing convolutional neural network for classifying gravitational waves from core-collapse supernovae
- Cloud-based XAI Services for Assessing Open Repository Models Under Adversarial Attacks
- CARLA: A Python Library to Benchmark Algorithmic Recourse and Counterfactual Explanation Algorithms
- VR-FuseNet: A Fusion of Heterogeneous Fundus Data and Explainable Deep Network for Diabetic Retinopathy Classification
- TAME: Attention Mechanism Based Feature Fusion for Generating Explanation Maps of Convolutional Neural Networks
- SSA: Semantic Structure Aware Inference for Weakly Pixel-Wise Dense Predictions without Cost
- Choose Your Explanation: A Comparison of SHAP and GradCAM in Human Activity Recognition
- Class Feature Pyramids for Video Explanation
- Beyond Fidelity: Explaining Vulnerability Localization of Learning-based Detectors
- Spatio-Temporal Perturbations for Video Attribution
- Representative Forgery Mining for Fake Face Detection
- LIMEcraft: Handcrafted superpixel selection and inspection for Visual eXplanations
- UM-CAM: Uncertainty-weighted Multi-resolution Class Activation Maps for Weakly-supervised Fetal Brain Segmentation
- T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
- Min-max Entropy for Weakly Supervised Pointwise Localization
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
- Exploring the Effect of Image Enhancement Techniques on COVID-19 Detection using Chest X-rays Images
- Hybrid Quantum Machine Learning Assisted Classification of COVID-19 from Computed Tomography Scans
- Adaptive Label Smoothing
- CounteRGAN: Generating Realistic Counterfactuals with Residual Generative Adversarial Nets
- Proactive Pseudo-Intervention: Causally Informed Contrastive Learning For Interpretable Vision Models
- On the Evaluation of the Plausibility and Faithfulness of Sentiment Analysis Explanations
- Multiple Sound Sources Localization from Coarse to Fine
- Classification Metrics for Image Explanations: Towards Building Reliable XAI-Evaluations
- Regularizing Reasons for Outfit Evaluation with Gradient Penalty
- A Challenging Benchmark of Anime Style Recognition
- Cross-modal Cognitive Consensus guided Audio-Visual Segmentation
- DBIA: Data-free Backdoor Injection Attack against Transformer Networks
- A Game-Theoretic Taxonomy of Visual Concepts in DNNs
- Probing the Purview of Neural Networks via Gradient Analysis
- Evaluation of Explanation Methods of AI -- CNNs in Image Classification Tasks with Reference-based and No-reference Metrics
- Transfer Learning and Explainable AI for Brain Tumor Classification: A Study Using MRI Data from Bangladesh
- What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain
- Study on the Helpfulness of Explainable Artificial Intelligence
- BSED: Baseline Shapley-Based Explainable Detector
- A Fine-Grained Vehicle Detection (FGVD) Dataset for Unconstrained Roads
- Robust Attentive Deep Neural Network for Exposing GAN-generated Faces
- Towards explainable artificial intelligence (XAI) for early anticipation of traffic accidents
- Contrastive Learning for Robust Android Malware Familial Classification
- CheXseg: Combining Expert Annotations with DNN-generated Saliency Maps for X-ray Segmentation
- Graph Representation learning for Audio & Music genre Classification
- Robust Models Are More Interpretable Because Attributions Look Normal
- ADVISE: ADaptive Feature Relevance and VISual Explanations for Convolutional Neural Networks
- Efficient Video Summarization Framework using EEG and Eye-tracking Signals
- Explainable Deep Learning Analysis for Raga Identification in Indian Art Music
- HealthiVert-GAN: A Novel Framework of Pseudo-Healthy Vertebral Image Synthesis for Interpretable Compression Fracture Grading
- CAMs as Shapley Value-based Explainers
- Memory Regulation and Alignment toward Generalizer RGB-Infrared Person
- Advancing Attribution-Based Neural Network Explainability through Relative Absolute Magnitude Layer-Wise Relevance Propagation and Multi-Component Evaluation
- An Explainable Model for EEG Seizure Detection based on Connectivity Features
- Harmonizing Feature Attributions Across Deep Learning Architectures: Enhancing Interpretability and Consistency
- Preserve, Promote, or Attack? GNN Explanation via Topology Perturbation
- Interpreting Interpretations: Organizing Attribution Methods by Criteria
- Learning Debiased and Disentangled Representations for Semantic Segmentation
- KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
- Towards Visually Explaining Similarity Models
- Better Understanding Differences in Attribution Methods via Systematic Evaluations
- Cross modal video representations for weakly supervised active speaker localization
- Deep Discriminative Representation Learning with Attention Map for Scene Classification
- FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
- SE3D: A Framework For Saliency Method Evaluation In 3D Imaging
- Optimizing machine learning methods to discover strong gravitational lenses in the Deep Lens Survey
- Characterizing out-of-distribution generalization of neural networks: application to the disordered Su-Schrieffer-Heeger model
- Understanding the Dependence of Perception Model Competency on Regions in an Image
- Zoom-CAM: Generating Fine-grained Pixel Annotations from Image Labels
- A Review of Explainable Artificial Intelligence in Manufacturing
- From Visual Explanations to Counterfactual Explanations with Latent Diffusion
- MamT: Multi-view Attention Networks for Mammography Cancer Classification
- Semantic Label Reduction Techniques for Autonomous Driving
- XDeep: An Interpretation Tool for Deep Neural Networks
- Gradient Weighted Superpixels for Interpretability in CNNs
- LID 2020: The Learning from Imperfect Data Challenge Results
- One-Vote Veto: Semi-Supervised Learning for Low-Shot Glaucoma Diagnosis
- Adapting Grad-CAM for Embedding Networks
- Towards Visually Explaining Video Understanding Networks with Perturbation
- PhagoStat a scalable and interpretable end to end framework for efficient quantification of cell phagocytosis in neurodegenerative disease studies
- MedicalPatchNet: A Patch-Based Self-Explainable AI Architecture for Chest X-ray Classification
- Adversarial TCAV -- Robust and Effective Interpretation of Intermediate Layers in Neural Networks
- SCANet: Split Coordinate Attention Network for Building Footprint Extraction
- Explainable Image Classification with Reduced Overconfidence for Tissue Characterisation
- POTHER: Patch-Voted Deep Learning-Based Chest X-ray Bias Analysis for COVID-19 Detection
- Pre or Post-Softmax Scores in Gradient-based Attribution Methods, What is Best?
- Can Perceptual Guidance Lead to Semantically Explainable Adversarial Perturbations?
- EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case
- Impact of data-splits on generalization: Identifying COVID-19 from cough and context
- Explaining neural network predictions of material strength
- DeepOpht: Medical Report Generation for Retinal Images via Deep Models and Visual Explanation
- Interpreting and Disentangling Feature Components of Various Complexity from DNNs
- Weakly Supervised Minirhizotron Image Segmentation with MIL-CAM
- Contextual Local Explanation for Black Box Classifiers
- A Step Towards Exposing Bias in Trained Deep Convolutional Neural Network Models
- How do Convolutional Neural Networks Learn Design?
- Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
- ViG-Bias: Visually Grounded Bias Discovery and Mitigation
- Integrated Grad-CAM: Sensitivity-Aware Visual Explanation of Deep Convolutional Networks via Integrated Gradient-Based Scoring
- Cluster Activation Mapping with Applications to Medical Imaging
- MFPP: Morphological Fragmental Perturbation Pyramid for Black-Box Model Explanations
- Gradient Frequency Modulation for Visually Explaining Video Understanding Models
- Meta-evaluating stability measures: MAX-Senstivity & AVG-Sensitivity
- Understanding of Kernels in CNN Models by Suppressing Irrelevant Visual Features in Images
- Towards Interpretable ANNs: An Exact Transformation to Multi-Class Multivariate Decision Trees
- Understanding Character Recognition using Visual Explanations Derived from the Human Visual System and Deep Networks
- CIM: Class-Irrelevant Mapping for Few-Shot Classification
- Sharpen Focus: Learning with Attention Separability and Consistency
- Deep Ensemble Collaborative Learning by using Knowledge-transfer Graph for Fine-grained Object Classification
- Improve CAM with Auto-adapted Segmentation and Co-supervised Augmentation
- Automated Cleanup of the ImageNet Dataset by Model Consensus, Explainability and Confident Learning
- Spatially-weighted Anomaly Detection with Regression Model
- Context-Gated Convolution
- Weakly-Supervised Cell Tracking via Backward-and-Forward Propagation
- Utilizing dataset affinity prediction in object detection to assess training data
- New Perspective of Interpretability of Deep Neural Networks
- Integrating Human Gaze into Attention for Egocentric Activity Recognition
- Visualizing Color-wise Saliency of Black-Box Image Classification Models
- Advancing Autonomous Driving: DepthSense with Radar and Spatial Attention
- Breast Cancer Classification in Deep Ultraviolet Fluorescence Images Using a Patch-Level Vision Transformer Framework
- Multi-scale discriminative Region Discovery for Weakly-Supervised Object Localization
- Soft Sensing Model Visualization: Fine-tuning Neural Network from What Model Learned
- Human-in-the-loop model explanation via verbatim boundary identification in generated neighborhoods
- A Method for Restoring the Training Set Distribution in an Image Classifier
- Towards Better Guided Attention and Human Knowledge Insertion in Deep Convolutional Neural Networks
- Robust Attacks on Deep Learning Face Recognition in the Physical World
- Interpretable by Design: Learning Predictors by Composing Interpretable Queries
- Where and When: Space-Time Attention for Audio-Visual Explanations
- Few Labels are all you need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series
- P-TAME: Explain Any Image Classifier with Trained Perturbations
- ExCon: Explanation-driven Supervised Contrastive Learning for Image Classification
- LaFAM: Unsupervised Feature Attribution with Label-free Activation Maps
- EDDA: Explanation-driven Data Augmentation to Improve Explanation Faithfulness
- A Peek Into the Reasoning of Neural Networks: Interpreting with Structural Visual Concepts
- Canonical Saliency Maps: Decoding Deep Face Models
- Learn to Rank: Visual Attribution by Learning Importance Ranking
- Metric-Guided Synthesis of Class Activation Mapping
- An Attention Self-supervised Contrastive Learning based Three-stage Model for Hand Shape Feature Representation in Cued Speech
- XPROAX-Local explanations for text classification with progressive neighborhood approximation
- A Survey of Machine Learning Techniques for Detecting and Diagnosing COVID-19 from Imaging
- Flip Learning: Erase to Segment
- Copy and Paste method based on Pose for Re-identification
- Transferring Knowledge with Attention Distillation for Multi-Domain Image-to-Image Translation
- Efficient Modelling Across Time of Human Actions and Interactions
- ProtoShotXAI: Using Prototypical Few-Shot Architecture for Explainable AI
- ST-ABN: Visual Explanation Taking into Account Spatio-temporal Information for Video Recognition
- Lost in Context: The Influence of Context on Feature Attribution Methods for Object Recognition
- Survival-oriented embeddings for improving accessibility to complex data structures
- Low-Cost Transfer Learning of Face Tasks
- Explainable, automated urban interventions to improve pedestrian and vehicle safety
- Defense Against Explanation Manipulation
- Defense-guided Transferable Adversarial Attacks
- MED-TEX: Transferring and Explaining Knowledge with Less Data from Pretrained Medical Imaging Models
- Play Fair: Frame Attributions in Video Models
- TorchPRISM: Principal Image Sections Mapping, a novel method for Convolutional Neural Network features visualization
- Right for the Right Reason: Making Image Classification Robust
- Ada-SISE: Adaptive Semantic Input Sampling for Efficient Explanation of Convolutional Neural Networks
- Human-Understandable Decision Making for Visual Recognition
- Beyond Occlusion: In Search for Near Real-Time Explainability of CNN-Based Prostate Cancer Classification
- How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
- Improving Attribution Methods by Learning Submodular Functions
- Weakly Supervised Instance Segmentation by Deep Community Learning
- Spatially Attentive Output Layer for Image Classification
- The Weighting Game: Evaluating Quality of Explainability Methods
- Learning Invariants through Soft Unification
- Saliency strikes back: How filtering out high frequencies improves white-box explanations
- SEEN: Sharpening Explanations for Graph Neural Networks using Explanations from Neighborhoods
- Visual Probing: Cognitive Framework for Explaining Self-Supervised Image Representations
- Indicative Image Retrieval: Turning Blackbox Learning into Grey
- Sanity Simulations for Saliency Methods
- How to Explain Neural Networks: an Approximation Perspective