The Landscape of Deep Learning Algorithms
arXiv:1705.07038
Abstract
This paper studies the landscape of empirical risk of deep neural networks by theoretically analyzing its convergence behavior to the population risk as well as its stationary points and properties. For an -layer linear neural network, we prove its empirical risk uniformly converges to its population risk at the rate of with training sample size of , the total weight dimension of and the magnitude bound of weight of each layer. We then derive the stability and generalization bounds for the empirical risk based on this result. Besides, we establish the uniform convergence of gradient of the empirical risk to its population counterpart. We prove the one-to-one correspondence of the non-degenerate stationary points between the empirical and population risks with convergence guarantees, which describes the landscape of deep neural networks. In addition, we analyze these properties for deep nonlinear neural networks with sigmoid activation functions. We prove similar results for convergence behavior of their empirical risks as well as the gradients and analyze properties of their non-degenerate stationary points. To our best knowledge, this work is the first one theoretically characterizing landscapes of deep learning algorithms. Besides, our results provide the sample complexity of training a good deep neural network. We also provide theoretical understanding on how the neural network depth , the layer width, the network size and parameter magnitude determine the neural network landscapes.
References in corpus (5)
Cited by in corpus (9)
- Learning Robust and High-Precision Quantum Controls
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced
- Stability and Generalization of Learning Algorithms that Converge to Global Optima
- The Multilinear Structure of ReLU Networks
- How Many Samples are Needed to Estimate a Convolutional or Recurrent Neural Network?
- Stationary Points of Shallow Neural Networks with Quadratic Activation Function
- Improved Learning of One-hidden-layer Convolutional Neural Networks with Overlaps
- Learning Graph Neural Networks with Approximate Gradient Descent