Detecting Statistical Interactions from Neural Network Weights
arXiv:1705.04977
Abstract
Interpreting neural networks is a crucial and challenging task in machine learning. In this paper, we develop a novel framework for detecting statistical interactions captured by a feedforward multilayer neural network by directly interpreting its learned weights. Depending on the desired interactions, our method can achieve significantly better or similar interaction detection performance compared to the state-of-the-art without searching an exponential solution space of possible interactions. We obtain this accuracy and efficiency by observing that interactions between input features are created by the non-additive effect of nonlinear activation functions, and that interacting paths are encoded in weight matrices. We demonstrate the performance of our method and the importance of discovered interactions via experimental results on both synthetic datasets and real-world application datasets.
Published in ICLR 2018
References in corpus (6)
- Towards A Rigorous Science of Interpretable Machine Learning
- Understanding Neural Networks Through Deep Visualization
- Predictive learning via rule ensembles
- Recurrent Models of Visual Attention
- Beyond Word Importance: Contextual Decomposition to Extract Interactions from LSTMs
- Interpretation of Prediction Models Using the Input Gradient
Cited by in corpus (14)
- Hierarchical interpretations for neural network predictions
- Explaining Explanations: Axiomatic Feature Interactions for Deep Networks
- Considerations When Learning Additive Explanations for Black-Box Models
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
- Enforcing Interpretability and its Statistical Impacts: Trade-offs between Accuracy and Interpretability
- Can I trust you more? Model-Agnostic Hierarchical Explanations
- Disentangled Attribution Curves for Interpreting Random Forests and Boosted Trees
- Hybrid Predictive Model: When an Interpretable Model Collaborates with a Black-box Model
- A Baseline for Shapley Values in MLPs: from Missingness to Neutrality
- Purifying Interaction Effects with the Functional ANOVA: An Efficient Algorithm for Recovering Identifiable Additive Models
- Mixture of Linear Models Co-supervised by Deep Neural Networks
- High Dimensional Model Explanations: an Axiomatic Approach
- Periodic Spectral Ergodicity: A Complexity Measure for Deep Neural Networks and Neural Architecture Search
- Relate and Predict: Structure-Aware Prediction with Jointly Optimized Neural DAG