Investigating the influence of noise and distractors on the interpretation of neural networks
arXiv:1611.07270
Abstract
Understanding neural networks is becoming increasingly important. Over the last few years different types of visualisation and explanation methods have been proposed. However, none of them explicitly considered the behaviour in the presence of noise and distracting elements. In this work, we will show how noise and distracting dimensions can influence the result of an explanation model. This gives a new theoretical insights to aid selection of the most appropriate explanation model within the deep-Taylor decomposition framework.
Presented at NIPS 2016 Workshop on Interpretable Machine Learning in Complex Systems
Cited by in corpus (17)
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- When Explanations Lie: Why Many Modified BP Attributions Fail
- On quantitative aspects of model interpretability
- TRUST-LAPSE: An Explainable and Actionable Mistrust Scoring Framework for Model Monitoring
- What went wrong and when? Instance-wise Feature Importance for Time-series Models
- Spatio-Temporal Perturbations for Video Attribution
- Order in the Court: Explainable AI Methods Prone to Disagreement
- Debugging Tests for Model Explanations
- Improving Feature Attribution through Input-specific Network Pruning
- CAMERAS: Enhanced Resolution And Sanity preserving Class Activation Mapping for image saliency
- A Simple Saliency Method That Passes the Sanity Checks
- Understanding and Diagnosing Vulnerability under Adversarial Attacks
- Learning Invariances for Interpretability using Supervised VAE
- DRR4Covid: Learning Automated COVID-19 Infection Segmentation from Digitally Reconstructed Radiographs
- It's FLAN time! Summing feature-wise latent representations for interpretability
- Deep Relevance Regularization: Interpretable and Robust Tumor Typing of Imaging Mass Spectrometry Data
- Attribution Mask: Filtering Out Irrelevant Features By Recursively Focusing Attention on Inputs of DNNs