The loss surface of deep and wide neural networks
arXiv:1704.08045
Abstract
While the optimization problem behind deep neural networks is highly non-convex, it is frequently observed in practice that training deep networks seems possible without getting stuck in suboptimal points. It has been argued that this is the case as all local minima are close to being globally optimal. We show that this is (almost) true, in fact almost all local minima are globally optimal, for a fully connected network with squared loss and analytic activation function given that the number of hidden units of one layer of the network is larger than the number of training points and the network structure from this layer on is pyramidal.
ICML 2017. Main results now hold for larger classes of loss functions
Cited by in corpus (90)
- Visualizing the Loss Landscape of Neural Nets
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Gradient Descent Finds Global Minima of Deep Neural Networks
- Essentially No Barriers in Neural Network Energy Landscape
- ResNet with one-neuron hidden layers is a Universal Approximator
- The Modern Mathematics of Deep Learning
- The Expressive Power of Neural Networks: A View from the Width
- A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks
- Optimal Approximation Rate of ReLU Networks in terms of Width and Depth
- Nonlinear Approximation via Compositions
- Depth with Nonlinearity Creates No Bad Local Minima in ResNets
- Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Global optimality conditions for deep neural networks
- On Interpretability of Artificial Neural Networks: A Survey
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- Effect of Depth and Width on Local Minima in Deep Learning
- Spurious Valleys in Two-layer Neural Network Optimization Landscapes
- Elimination of All Bad Local Minima in Deep Learning
- Implicit Bias of Gradient Descent on Linear Convolutional Networks
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- On the Decision Boundary of Deep Neural Networks
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
- Mathematical Models of Overparameterized Neural Networks
- Theoretical properties of the global optimizer of two layer neural network
- Characterization of Gradient Dominance and Regularity Conditions for Neural Networks
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- High Dimensional Spaces, Deep Learning and Adversarial Examples
- Data efficiency and extrapolation trends in neural network interatomic potentials
- LCA: Loss Change Allocation for Neural Network Training
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Neural Networks Should Be Wide Enough to Learn Disconnected Decision Regions
- Deep Network Approximation: Achieving Arbitrary Accuracy with Fixed Number of Neurons
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation Functions
- The Landscape of Deep Learning Algorithms
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- The Multilinear Structure of ReLU Networks
- Self-Paced Contrastive Learning for Semi-supervised Medical Image Segmentation with Meta-labels
- How Many Samples are Needed to Estimate a Convolutional or Recurrent Neural Network?
- Towards Robust Deep Neural Networks
- Stationary Points of Shallow Neural Networks with Quadratic Activation Function
- On Connected Sublevel Sets in Deep Learning
- Non-attracting Regions of Local Minima in Deep and Wide Neural Networks
- Beyond Random Matrix Theory for Deep Networks
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural Networks
- Noether: The More Things Change, the More Stay the Same
- Are deep ResNets provably better than linear predictors?
- Why Learning of Large-Scale Neural Networks Behaves Like Convex Optimization
- A Differential Topological View of Challenges in Learning with Feedforward Neural Networks
- Improved Learning of One-hidden-layer Convolutional Neural Networks with Overlaps
- Guaranteed Recovery of One-Hidden-Layer Neural Networks via Cross Entropy
- Avoiding Spurious Local Minima in Deep Quadratic Networks
- Translating Diffusion, Wavelets, and Regularisation into Residual Networks
- On the Universality of the Double Descent Peak in Ridgeless Regression
- Learning Dynamics of Linear Denoising Autoencoders
- Towards an Understanding of Residual Networks Using Neural Tangent Hierarchy (NTH)
- Abstraction Mechanisms Predict Generalization in Deep Neural Networks
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology
- The global optimum of shallow neural network is attained by ridgelet transform
- Collective evolution of weights in wide neural networks
- On the Global Convergence of Gradient Descent for multi-layer ResNets in the mean-field regime
- On the achievability of blind source separation for high-dimensional nonlinear source mixtures
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension
- The loss landscape of deep linear neural networks: a second-order analysis
- Ridge Regression with Over-Parametrized Two-Layer Networks Converge to Ridgelet Spectrum
- Insights into Ordinal Embedding Algorithms: A Systematic Evaluation
- Measure, Manifold, Learning, and Optimization: A Theory Of Neural Networks
- A Generative Neural Network Framework for Automated Software Testing
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU Networks
- Neural Networks with Complex-Valued Weights Have No Spurious Local Minima
- Fuzzy Logic Interpretation of Quadratic Networks
- The layer-wise L1 Loss Landscape of Neural Nets is more complex around local minima
- Information-Theoretic Perspective of Federated Learning
- Theoretical Exploration of Flexible Transmitter Model
- Achieving Small Test Error in Mildly Overparameterized Neural Networks
- Analysis and Optimisation of Bellman Residual Errors with Neural Function Approximation
- Deep Neural Networks Are Congestion Games: From Loss Landscape to Wardrop Equilibrium and Beyond
- Ghosts in Neural Networks: Existence, Structure and Role of Infinite-Dimensional Null Space
- Understanding How Over-Parametrization Leads to Acceleration: A case of learning a single teacher neuron
- When Can Neural Networks Learn Connected Decision Regions?
- When Are Solutions Connected in Deep Networks?
- CNNs are Globally Optimal Given Multi-Layer Support
- Benefits of over-parameterization with EM
- Training Deep Neural Networks via Branch-and-Bound
- On the alpha-loss Landscape in the Logistic Model
- The Hidden Convex Optimization Landscape of Two-Layer ReLU Neural Networks: an Exact Characterization of the Optimal Solutions