The Global Landscape of Neural Networks: An Overview
arXiv:2007.01429 · doi:10.1109/MSP.2020.3004124
Abstract
One of the major concerns for neural network training is that the non-convexity of the associated loss functions may cause bad landscape. The recent success of neural networks suggests that their loss landscape is not too bad, but what specific results do we know about the landscape? In this article, we review recent findings and results on the global landscape of neural networks. First, we point out that wide neural nets may have sub-optimal local minima under certain assumptions. Second, we discuss a few rigorous results on the geometric properties of wide networks such as "no bad basin", and some modifications that eliminate sub-optimal local minima and/or decreasing paths to infinity. Third, we discuss visualization and empirical explorations of the landscape for practical neural nets. Finally, we briefly discuss some convergence results and their relation to landscape results.
16 pages. 8 figures
References in corpus (6)
- The Loss Surfaces of Multilayer Networks
- How to Escape Saddle Points Efficiently
- Qualitatively characterizing neural network optimization problems
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- Entropic gradient descent algorithms and wide flat minima
- Revisiting Landscape Analysis in Deep Neural Networks: Eliminating Decreasing Paths to Infinity
Cited by in corpus (13)
- Deep matrix factorizations
- A Geometric Analysis of Neural Collapse with Unconstrained Features
- Mathematical Models of Overparameterized Neural Networks
- Understanding and Improving Model Averaging in Federated Learning on Heterogeneous Data
- Taxonomizing local versus global structure in neural network loss landscapes
- Training a Two Layer ReLU Network Analytically
- AdapMTL: Adaptive Pruning Framework for Multitask Learning Model
- The loss landscape of deep linear neural networks: a second-order analysis
- Non-Convex Exact Community Recovery in Stochastic Block Model
- Neural Networks with Complex-Valued Weights Have No Spurious Local Minima
- On the Stability Properties and the Optimization Landscape of Training Problems with Squared Loss for Neural Networks and General Nonlinear Conic Approximation Schemes
- WGAN with an Infinitely Wide Generator Has No Spurious Stationary Points
- Embedding Principle: a hierarchical structure of loss landscape of deep neural networks