Theory II: Landscape of the Empirical Risk in Deep Learning
arXiv:1703.09833
Abstract
Previous theoretical work on deep learning and neural network optimization tend to focus on avoiding saddle points and local minima. However, the practical observation is that, at least in the case of the most successful Deep Convolutional Neural Networks (DCNNs), practitioners can always increase the network size to fit the training data (an extreme example would be [1]). The most successful DCNNs such as VGG and ResNets are best used with a degree of "overparametrization". In this work, we characterize with a mix of theory and experiments, the landscape of the empirical risk of overparametrized DCNNs. We first prove in the regression framework the existence of a large number of degenerate global minimizers with zero empirical error (modulo inconsistent equations). The argument that relies on the use of Bezout theorem is rigorous when the RELUs are replaced by a polynomial nonlinearity (which empirically works as well). As described in our Theory III [2] paper, the same minimizers are degenerate and thus very likely to be found by SGD that will furthermore select with higher probability the most robust zero-minimizer. We further experimentally explored and visualized the landscape of empirical risk of a DCNN on CIFAR-10 during the entire training process and especially the global minima. Finally, based on our theoretical and experimental results, we propose an intuitive model of the landscape of DCNN's empirical loss surface, which might not be as complicated as people commonly believe.
Merged figures to make the main text more compact. Moved some similar figures to the appendix
References in corpus (1)
Cited by in corpus (26)
- Visualizing the Loss Landscape of Neural Nets
- Neural network models and deep learning - a primer for biologists
- Nonparametric regression using deep neural networks with ReLU activation function
- Optimization for deep learning: theory and algorithms
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- Theory of Deep Learning III: explaining the non-overfitting puzzle
- Theory of Deep Learning IIb: Optimization Properties of SGD
- On the Decision Boundary of Deep Neural Networks
- High Dimensional Spaces, Deep Learning and Adversarial Examples
- A Surprising Linear Relationship Predicts Test Performance in Deep Networks
- Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization
- Deep Hamiltonian networks based on symplectic integrators
- Sparse Deep Neural Network Exact Solutions
- Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness
- Theory IIIb: Generalization in Deep Networks
- Non-Gaussian processes and neural networks at finite widths
- Theory III: Dynamics and Generalization in Deep Networks
- McKernel: A Library for Approximate Kernel Expansions in Log-linear Time
- Understanding Global Loss Landscape of One-hidden-layer ReLU Networks, Part 1: Theory
- Understanding Global Loss Landscape of One-hidden-layer ReLU Networks, Part 2: Experiments and Analysis
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations
- Towards an understanding of CNNs: analysing the recovery of activation pathways via Deep Convolutional Sparse Coding
- When Can Neural Networks Learn Connected Decision Regions?
- Layer Dynamics of Linearised Neural Nets
- Theory of Generative Deep Learning : Probe Landscape of Empirical Error via Norm Based Capacity Control