Landscape analysis for shallow neural networks: complete classification of critical points for affine target functions
arXiv:2103.10922 · doi:10.1007/s00332-022-09823-8
Abstract
In this paper, we analyze the landscape of the true loss of neural networks with one hidden layer and ReLU, leaky ReLU, or quadratic activation. In all three cases, we provide a complete classification of the critical points in the case where the target function is affine and one-dimensional. In particular, we show that there exist no local maxima and clarify the structure of saddle points. Moreover, we prove that non-global local minima can only be caused by `dead' ReLU neurons. In particular, they do not appear in the case of leaky ReLU or quadratic activation. Our approach is of a combinatorial nature and builds on a careful analysis of the different types of hidden neurons that can occur.
References in corpus (1)
Cited by in corpus (7)
- A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
- Convergence analysis for gradient flows in the training of artificial neural networks with ReLU activation
- On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks
- A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
- Existence, uniqueness, and convergence rates for gradient flows in the training of artificial neural networks with ReLU activation
- Gradient descent provably escapes saddle points in the training of shallow ReLU networks
- Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks