Critical Points of Neural Networks: Analytical Forms and Landscape Properties
arXiv:1710.11205
Abstract
Due to the success of deep learning to solving a variety of challenging machine learning tasks, there is a rising interest in understanding loss functions for training neural networks from a theoretical aspect. Particularly, the properties of critical points and the landscape around them are of importance to determine the convergence performance of optimization algorithms. In this paper, we provide full (necessary and sufficient) characterization of the analytical forms for the critical points (as well as global minimizers) of the square loss functions for various neural networks. We show that the analytical forms of the critical points characterize the values of the corresponding loss functions as well as the necessary and sufficient conditions to achieve global minimum. Furthermore, we exploit the analytical forms of the critical points to characterize the landscape properties for the loss functions of these neural networks. One particular conclusion is that: The loss function of linear networks has no spurious local minimum, while the loss function of one-hidden-layer nonlinear networks with ReLU activation function does have local minimum that is not global minimum.
References in corpus (1)
Cited by in corpus (11)
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Optimization for deep learning: theory and algorithms
- Mathematical Models of Overparameterized Neural Networks
- Understanding Deep Learning via Decision Boundary
- Over Parameterized Two-level Neural Networks Can Learn Near Optimal Feature Representations
- The Local Elasticity of Neural Networks
- On Connected Sublevel Sets in Deep Learning
- Numerically Recovering the Critical Points of a Deep Linear Autoencoder
- On the Stability Properties and the Optimization Landscape of Training Problems with Squared Loss for Neural Networks and General Nonlinear Conic Approximation Schemes
- The Landscape of Multi-Layer Linear Neural Network From the Perspective of Algebraic Geometry
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations