On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
arXiv:1812.11039
Abstract
Wide networks are often believed to have a nice optimization landscape, but what rigorous results can we prove? To understand the benefit of width, it is important to identify the difference between wide and narrow networks. In this work, we prove that from narrow to wide networks, there is a phase transition from having sub-optimal basins to no sub-optimal basins. Specifically, we prove two results: on the positive side, for any continuous activation functions, the loss surface of a class of wide networks has no sub-optimal basins, where "basin" is defined as the set-wise strict local minimum; on the negative side, for a large class of networks with width below a threshold, we construct strict local minima that are not global. These two results together show the phase transition from narrow to wide networks.
ver1: Nov 22, 2018; ver2: Jan 20, 2020; ver3: July 26, 2020; ver4: Jan 19, 2021; ver5: Sept 2, 2021
References in corpus (32)
- Large Scale GAN Training for High Fidelity Natural Image Synthesis
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- The Loss Surfaces of Multilayer Networks
- A Convergence Theory for Deep Learning via Over-Parameterization
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
- The generalization error of random features regression: Precise asymptotics and double descent curve
- How to Escape Saddle Points Efficiently
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Spectrally-normalized margin bounds for neural networks
- On the Power of Over-parametrization in Neural Networks with Quadratic Activation
- Learning ReLU Networks on Linearly Separable Data: Algorithm, Optimality, and Generalization
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- Spurious Local Minima are Common in Two-Layer ReLU Neural Networks
- Recovery Guarantees for One-hidden-layer Neural Networks
- The jamming transition as a paradigm to understand the loss landscape of deep neural networks
- On the Computational Efficiency of Training Neural Networks
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks
- When is a Convolutional Filter Easy To Learn?
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Accelerated Gradient Descent Escapes Saddle Points Faster than Gradient Descent
- Understanding the Loss Surface of Neural Networks for Binary Classification
- Spurious Valleys in Two-layer Neural Network Optimization Landscapes
- A mean-field limit for certain deep neural networks
- Are ResNets Provably Better than Linear Predictors?
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- The Difficulty of Training Sparse Neural Networks
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
- Learning Neural Networks with Two Nonlinear Layers in Polynomial Time
- Theoretical properties of the global optimizer of two layer neural network
- A theory on the absence of spurious solutions for nonconvex and nonsmooth optimization
- On Connected Sublevel Sets in Deep Learning
Cited by in corpus (15)
- Learning ReLU Networks on Linearly Separable Data: Algorithm, Optimality, and Generalization
- Optimization for deep learning: theory and algorithms
- The Global Landscape of Neural Networks: An Overview
- Revisiting Landscape Analysis in Deep Neural Networks: Eliminating Decreasing Paths to Infinity
- Dynamic of Stochastic Gradient Descent with State-Dependent Noise
- Understanding Global Loss Landscape of One-hidden-layer ReLU Networks, Part 1: Theory
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology
- Global Convergence and Generalization Bound of Gradient-Based Meta-Learning with Deep Neural Nets
- Optimization Landscapes of Wide Deep Neural Networks Are Benign
- Neural Networks with Complex-Valued Weights Have No Spurious Local Minima
- Effect of the initial configuration of weights on the training and function of artificial neural networks
- On the Stability Properties and the Optimization Landscape of Training Problems with Squared Loss for Neural Networks and General Nonlinear Conic Approximation Schemes
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations
- WGAN with an Infinitely Wide Generator Has No Spurious Stationary Points
- When Are Solutions Connected in Deep Networks?