Faster SGD Using Sketched Conditioning
arXiv:1506.02649
Abstract
We propose a novel method for speeding up stochastic optimization algorithms via sketching methods, which recently became a powerful tool for accelerating algorithms for numerical linear algebra. We revisit the method of conditioning for accelerating first-order methods and suggest the use of sketching methods for constructing a cheap conditioner that attains a significant speedup with respect to the Stochastic Gradient Descent (SGD) algorithm. While our theoretical guarantees assume convexity, we discuss the applicability of our method to deep neural networks, and experimentally demonstrate its merits.
References in corpus (1)
Cited by in corpus (6)
- Shampoo: Preconditioned Stochastic Tensor Optimization
- Efficient Second Order Online Learning by Sketching
- Scalable Second Order Optimization for Deep Learning
- Matrix-Free Preconditioning in Online Learning
- From Persistent Homology to Reinforcement Learning with Applications for Retail Banking
- Optimal Sketching Bounds for Exp-concave Stochastic Minimization