Preconditioned Stochastic Gradient Descent
arXiv:1512.04202 · doi:10.1109/TNNLS.2017.2672978
Abstract
Stochastic gradient descent (SGD) still is the workhorse for many practical problems. However, it converges slow, and can be difficult to tune. It is possible to precondition SGD to accelerate its convergence remarkably. But many attempts in this direction either aim at solving specialized problems, or result in significantly more complicated methods than SGD. This paper proposes a new method to estimate a preconditioner such that the amplitudes of perturbations of preconditioned stochastic gradient match that of the perturbations of parameters to be optimized in a way comparable to Newton method for deterministic optimization. Unlike the preconditioners based on secant equation fitting as done in deterministic quasi-Newton methods, which assume positive definite Hessian and approximate its inverse, the new preconditioner works equally well for both convex and non-convex optimizations with exact or noisy gradients. When stochastic gradient is used, it can naturally damp the gradient noise to stabilize SGD. Efficient preconditioner estimation methods are developed, and with reasonable simplifications, they are applicable to large scaled problems. Experimental results demonstrate that equipped with the new preconditioner, without any tuning effort, preconditioned SGD can efficiently solve many challenging problems like the training of a deep neural network or a recurrent neural network requiring extremely long term memories.
13 pages, 9 figures. To appear in IEEE Transactions on Neural Networks and Learning Systems. Supplemental materials on https://sites.google.com/site/lixilinx/home/psgd
References in corpus (3)
Cited by in corpus (17)
- SoftAdapt: Techniques for Adaptive Loss Weighting of Neural Networks with Multi-Part Loss Functions
- The Strength of Nesterov's Extrapolation in the Individual Convergence of Nonsmooth Optimization
- A variable metric mini-batch proximal stochastic recursive gradient algorithm with diagonal Barzilai-Borwein stepsize
- Distributed Machine Learning for Wireless Communication Networks: Techniques, Architectures, and Applications
- Deep Particulate Matter Forecasting Model Using Correntropy-Induced Loss
- Accelerating Mini-batch SARAH by Step Size Rules
- Independent Vector Analysis with Deep Neural Network Source Priors
- Understanding Short-Range Memory Effects in Deep Neural Networks
- Error Bounds and Applications for Stochastic Approximation with Non-Decaying Gain
- First-Order Preconditioning via Hypergradient Descent
- Efficient Implementation of Second-Order Stochastic Approximation Algorithms in High-Dimensional Problems
- Recurrent neural network training with preconditioned stochastic gradient descent
- Efficient preconditioned stochastic gradient descent for estimation in latent variable models
- Hybrid Acceleration Scheme for Variance Reduced Stochastic Optimization Algorithms
- Online Second Order Methods for Non-Convex Stochastic Optimizations
- Gaussian Mean Field Regularizes by Limiting Learned Information
- Preconditioner on Matrix Lie Group for SGD