Certifiably Robust Interpretation in Deep Learning
arXiv:1905.12105
Abstract
Deep learning interpretation is essential to explain the reasoning behind model predictions. Understanding the robustness of interpretation methods is important especially in sensitive domains such as medical applications since interpretation results are often used in downstream tasks. Although gradient-based saliency maps are popular methods for deep learning interpretation, recent works show that they can be vulnerable to adversarial attacks. In this paper, we address this problem and provide a certifiable defense method for deep learning interpretation. We show that a sparsified version of the popular SmoothGrad method, which computes the average saliency maps over random perturbations of the input, is certifiably robust against adversarial perturbations. We obtain this result by extending recent bounds for certifiably robust smooth classifiers to the interpretation setting. Experiments on ImageNet samples validate our theory.
References in corpus (8)
- SmoothGrad: removing noise by adding noise
- Sanity Checks for Saliency Maps
- Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers
- Certified Adversarial Robustness with Additive Noise
- Decoupled Deep Neural Network for Semi-supervised Semantic Segmentation
- On the Effectiveness of Defensive Distillation
- Towards Robust, Locally Linear Deep Networks
- Robust Saliency Detection via Fusing Foreground and Background Priors
Cited by in corpus (11)
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- (De)Randomized Smoothing for Certifiable Defense against Patch Attacks
- Are Perceptually-Aligned Gradients a General Property of Robust Classifiers?
- Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
- Improved, Deterministic Smoothing for L_1 Certified Robustness
- Adversarial Attacks and Defenses: An Interpretation Perspective
- Towards the Unification and Robustness of Perturbation and Gradient Based Explanations
- Tight Second-Order Certificates for Randomized Smoothing
- DANCE: Enhancing saliency maps using decoys
- Deep Class-Specific Affinity-Guided Convolutional Network for Multimodal Unpaired Image Segmentation
- Defense Against Explanation Manipulation