Optimization for deep learning: theory and algorithms
arXiv:1912.08957
Abstract
When and why can a neural network be successfully trained? This article provides an overview of optimization algorithms and theory for training neural networks. First, we discuss the issue of gradient explosion/vanishing and the more general issue of undesirable spectrum, and then discuss practical solutions including careful initialization and normalization methods. Second, we review generic optimization methods used in training neural networks, such as SGD, adaptive gradient methods and distributed methods, and theoretical results for these algorithms. Third, we review existing research on the global issues of neural network training, including results on bad local minima, mode connectivity, lottery ticket hypothesis and infinite-width analysis.
38 pages of main body; 5 pages of appendix; 12 pages of references
References in corpus (31)
- ADADELTA: An Adaptive Learning Rate Method
- Neural Architecture Search with Reinforcement Learning
- On the Convergence of Adam and Beyond
- Understanding deep learning requires rethinking generalization
- The Loss Surfaces of Multilayer Networks
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Improving Generalization Performance by Switching from Adam to SGD
- No More Pesky Learning Rates
- Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Qualitatively characterizing neural network optimization problems
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
- Learning One-hidden-layer Neural Networks with Landscape Design
- Fixup Initialization: Residual Learning Without Normalization
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Mathematics of Deep Learning
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- Yet Another Accelerated SGD: ResNet-50 Training on ImageNet in 74.7 seconds
- A mean-field limit for certain deep neural networks
- Gradient Dynamics of Shallow Univariate ReLU Networks
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- Theoretical properties of the global optimizer of two layer neural network
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Dynamical Isometry and a Mean Field Theory of LSTMs and GRUs
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- Positively Scale-Invariant Flatness of ReLU Neural Networks
- On Connected Sublevel Sets in Deep Learning