Deep Learning Through the Lens of Example Difficulty
arXiv:2106.09647
Abstract
Existing work on understanding deep learning often employs measures that compress all data-dependent information into a few numbers. In this work, we adopt a perspective based on the role of individual examples. We introduce a measure of the computational difficulty of making a prediction for a given input: the (effective) prediction depth. Our extensive investigation reveals surprising yet simple relationships between the prediction depth of a given input and the model's uncertainty, confidence, accuracy and speed of learning for that data point. We further categorize difficult examples into three interpretable groups, demonstrate how these groups are processed differently inside deep models and showcase how this understanding allows us to improve prediction accuracy. Insights from our study lead to a coherent view of a number of separately reported phenomena in the literature: early layers generalize while later layers memorize; early layers converge faster and networks learn easy data and simple functions first.
Main paper: 15 pages, 8 figures. Appendix: 31 pages, 40 figures
References in corpus (8)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- BatchEnsemble: An Alternative Approach to Efficient Ensemble and Lifelong Learning
- On the Origin of Implicit Regularization in Stochastic Gradient Descent
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
- Distribution Density, Tails, and Outliers in Machine Learning: Metrics and Applications
- On the geometry of generalization and memorization in deep neural networks