Exponentially vanishing sub-optimal local minima in multilayer neural networks
arXiv:1702.05777
Abstract
Background: Statistical mechanics results (Dauphin et al. (2014); Choromanska et al. (2015)) suggest that local minima with high error are exponentially rare in high dimensions. However, to prove low error guarantees for Multilayer Neural Networks (MNNs), previous works so far required either a heavily modified MNN model or training method, strong assumptions on the labels (e.g., "near" linear separability), or an unrealistic hidden layer with units. Results: We examine a MNN with one hidden layer of piecewise linear units, a single output, and a quadratic loss. We prove that, with high probability in the limit of datapoints, the volume of differentiable regions of the empiric loss containing sub-optimal differentiable local minima is exponentially vanishing in comparison with the same volume of global minima, given standard normal input of dimension , and a more realistic number of hidden units. We demonstrate our results numerically: for example, binary classification training error on CIFAR with only hidden neurons.
Cited by in corpus (26)
- Visualizing the Loss Landscape of Neural Nets
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Depth Creates No Bad Local Minima
- Depth with Nonlinearity Creates No Bad Local Minima in ResNets
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- Understanding the Loss Surface of Neural Networks for Binary Classification
- Elimination of All Bad Local Minima in Deep Learning
- On the Decision Boundary of Deep Neural Networks
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- On the Power and Limitations of Random Features for Understanding Neural Networks
- Recent advances in deep learning theory
- The Landscape of Deep Learning Algorithms
- Generalization in Machine Learning via Analytical Learning Theory
- Fix your classifier: the marginal value of training the last weight layer
- The Multilinear Structure of ReLU Networks
- The Local Elasticity of Neural Networks
- Trusted Artificial Intelligence: Towards Certification of Machine Learning Applications
- Non-attracting Regions of Local Minima in Deep and Wide Neural Networks
- Learning Dynamics of Linear Denoising Autoencoders
- The Global Optimization Geometry of Shallow Linear Neural Networks
- Neural networks behave as hash encoders: An empirical study
- Deep Inverse Optimization
- Deep Neural Networks Are Congestion Games: From Loss Landscape to Wardrop Equilibrium and Beyond
- When Are Solutions Connected in Deep Networks?