Noise Modulation: Let Your Model Interpret Itself
arXiv:2103.10603
Abstract
Given the great success of Deep Neural Networks(DNNs) and the black-box nature of it,the interpretability of these models becomes an important issue.The majority of previous research works on the post-hoc interpretation of a trained model.But recently, adversarial training shows that it is possible for a model to have an interpretable input-gradient through training.However,adversarial training lacks efficiency for interpretability.To resolve this problem, we construct an approximation of the adversarial perturbations and discover a connection between adversarial training and amplitude modulation. Based on a digital analogy,we propose noise modulation as an efficient and model-agnostic alternative to train a model that interprets itself with input-gradients.Experiment results show that noise modulation can effectively increase the interpretability of input-gradients model-agnosticly.
References in corpus (8)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- SmoothGrad: removing noise by adding noise
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Interpreting Adversarially Trained Convolutional Neural Networks
- On the Connection Between Adversarial Robustness and Saliency Map Interpretability
- Bridging Adversarial Robustness and Gradient Interpretability
- A Survey on Deep Learning Methods for Semantic Image Segmentation in Real-Time