Small nonlinearities in activation functions create bad local minima in neural networks
arXiv:1802.03487
Abstract
We investigate the loss surface of neural networks. We prove that even for one-hidden-layer networks with "slightest" nonlinearity, the empirical risks have spurious local minima in most cases. Our results thus indicate that in general "no spurious local minima" is a property limited to deep linear networks, and insights obtained from linear networks may not be robust. Specifically, for ReLU(-like) networks we constructively prove that for almost all practical datasets there exist infinitely many local minima. We also present a counterexample for more general activations (sigmoid, tanh, arctan, ReLU, etc.), for which there exists a bad local minimum. Our results make the least restrictive assumptions relative to existing results on spurious local optima in neural networks. We complete our discussion by presenting a comprehensive characterization of global optimality for deep linear networks, which unifies other results on this topic.
33 pages, appeared at ICLR 2019
Cited by in corpus (18)
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- A Geometric Analysis of Neural Collapse with Unconstrained Features
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- The Usual Suspects? Reassessing Blame for VAE Posterior Collapse
- Recent advances in deep learning theory
- Depth creates no more spurious local minima
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Tighter Generalization Bounds for Iterative Differentially Private Learning Algorithms
- Are deep ResNets provably better than linear predictors?
- Understanding Global Loss Landscape of One-hidden-layer ReLU Networks, Part 1: Theory
- A Note on Connectivity of Sublevel Sets in Deep Learning
- Training Two-Layer ReLU Networks with Gradient Descent is Inconsistent
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension
- Neural Networks with Complex-Valued Weights Have No Spurious Local Minima
- When Are Solutions Connected in Deep Networks?
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations
- The Expressivity and Training of Deep Neural Networks: toward the Edge of Chaos?