The Nonlinearity Coefficient - Predicting Generalization in Deep Neural Networks
arXiv:1806.00179
Abstract
For a long time, designing neural architectures that exhibit high performance was considered a dark art that required expert hand-tuning. One of the few well-known guidelines for architecture design is the avoidance of exploding gradients, though even this guideline has remained relatively vague and circumstantial. We introduce the nonlinearity coefficient (NLC), a measurement of the complexity of the function computed by a neural network that is based on the magnitude of the gradient. Via an extensive empirical study, we show that the NLC is a powerful predictor of test error and that attaining a right-sized NLC is essential for optimal performance. The NLC exhibits a range of intriguing and important properties. It is closely tied to the amount of information gained from computing a single network gradient. It is tied to the error incurred when replacing the nonlinearity operations in the network with linear operations. It is not susceptible to the confounders of multiplicative scaling, additive bias and layer width. It is stable from layer to layer. Hence, we argue that the NLC is the first robust predictor of overfitting in deep networks.
Previous name: The Nonlinearity Coefficient - Predicting Overfitting in Deep Neural Networks
References in corpus (13)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- On the difficulty of training Recurrent Neural Networks
- Self-Normalizing Neural Networks
- Layer Normalization
- Exponential expressivity in deep neural networks through transient chaos
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Predicting the Generalization Gap in Deep Networks with Margin Distributions
- The Shattered Gradients Problem: If resnets are the answer, then what is the question?
- Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks
- A Mean Field Theory of Batch Normalization
- Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10,000-Layer Vanilla Convolutional Neural Networks
- The exploding gradient problem demystified - definition, prevalence, impact, origin, tradeoffs, and solutions
- Deep Information Propagation