HYDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural Networks
arXiv:2102.02515 · doi:10.1609/aaai.v35i8.16871
Abstract
The behaviors of deep neural networks (DNNs) are notoriously resistant to human interpretations. In this paper, we propose Hypergradient Data Relevance Analysis, or HYDRA, which interprets the predictions made by DNNs as effects of their training data. Existing approaches generally estimate data contributions around the final model parameters and ignore how the training data shape the optimization trajectory. By unrolling the hypergradient of test loss w.r.t. the weights of training data, HYDRA assesses the contribution of training data toward test data points throughout the training trajectory. In order to accelerate computation, we remove the Hessian from the calculation and prove that, under moderate conditions, the approximation error is bounded. Corroborating this theoretical claim, empirical results indicate the error is indeed small. In addition, we quantitatively demonstrate that HYDRA outperforms influence functions in accurately estimating data contribution and detecting noisy data labels. The source code is available at https://github.com/cyyever/aaai_hydra_8686.
References in corpus (29)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Axiomatic Attribution for Deep Networks
- Striving for Simplicity: The All Convolutional Net
- Understanding Black-box Predictions via Influence Functions
- This Looks Like That: Deep Learning for Interpretable Image Recognition
- Bird Species Categorization Using Pose Normalized Deep Convolutional Nets
- Deep Learning is Robust to Massive Label Noise
- Gradient-based Hyperparameter Optimization through Reversible Learning
- Meta-Weight-Net: Learning an Explicit Mapping For Sample Weighting
- FACE: Feasible and Actionable Counterfactual Explanations
- Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
- Towards Efficient Data Valuation Based on the Shapley Value
- Data Shapley: Equitable Valuation of Data for Machine Learning
- On the Accuracy of Influence Functions for Measuring Group Effects
- Representer Point Selection for Explaining Deep Neural Networks
- Data Valuation using Reinforcement Learning
- Understanding and correcting pathologies in the training of learned optimizers
- Optimizing Millions of Hyperparameters by Implicit Differentiation
- NormLime: A New Feature Importance Metric for Explaining Deep Neural Networks
- Interpretable & Explorable Approximations of Black Box Models
- Data Cleansing for Models Trained with SGD
- Penalty Method for Inversion-Free Deep Bilevel Optimization
- Approximation Trees: Statistical Stability in Model Distillation
- Learning Norms from Stories: A Prior for Value Aligned Agents
- RelatIF: Identifying Explanatory Training Examples via Relative Influence
- SOSELETO: A Unified Approach to Transfer Learning and Training with Noisy Labels
- Frequentist Uncertainty in Recurrent Neural Networks via Blockwise Influence Functions
- Explaining Neural Networks Semantically and Quantitatively
- Multi-Stage Influence Function
Cited by in corpus (5)
- Training Data Influence Analysis and Estimation: A Survey
- Better, Not Just More: Data-Centric Machine Learning for Earth Observation
- Identifying a Training-Set Attack's Target Using Renormalized Influence Estimation
- Rethinking Influence Functions of Neural Networks in the Over-parameterized Regime
- Data Cleansing for GANs