Investigating the influence of noise and distractors on the interpretation of neural networks
arXiv:1611.07270
Abstract
Understanding neural networks is becoming increasingly important. Over the last few years different types of visualisation and explanation methods have been proposed. However, none of them explicitly considered the behaviour in the presence of noise and distracting elements. In this work, we will show how noise and distracting dimensions can influence the result of an explanation model. This gives a new theoretical insights to aid selection of the most appropriate explanation model within the deep-Taylor decomposition framework.
Presented at NIPS 2016 Workshop on Interpretable Machine Learning in Complex Systems
Cited by in corpus (11)
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- On quantitative aspects of model interpretability
- Spatio-Temporal Perturbations for Video Attribution
- Debugging Tests for Model Explanations
- CAMERAS: Enhanced Resolution And Sanity preserving Class Activation Mapping for image saliency
- A Simple Saliency Method That Passes the Sanity Checks
- Understanding and Diagnosing Vulnerability under Adversarial Attacks
- Learning Invariances for Interpretability using Supervised VAE
- DRR4Covid: Learning Automated COVID-19 Infection Segmentation from Digitally Reconstructed Radiographs
- Attribution Mask: Filtering Out Irrelevant Features By Recursively Focusing Attention on Inputs of DNNs
- Deep Relevance Regularization: Interpretable and Robust Tumor Typing of Imaging Mass Spectrometry Data