Fast Stochastic Variance Reduced Gradient Method with Momentum Acceleration for Machine Learning
arXiv:1703.07948
Abstract
Recently, research on accelerated stochastic gradient descent methods (e.g., SVRG) has made exciting progress (e.g., linear convergence for strongly convex problems). However, the best-known methods (e.g., Katyusha) requires at least two auxiliary variables and two momentum parameters. In this paper, we propose a fast stochastic variance reduction gradient (FSVRG) method, in which we design a novel update rule with the Nesterov's momentum and incorporate the technique of growing epoch size. FSVRG has only one auxiliary variable and one momentum weight, and thus it is much simpler and has much lower per-iteration complexity. We prove that FSVRG achieves linear convergence for strongly convex problems and the optimal convergence rate for non-strongly convex problems, where is the number of outer-iterations. We also extend FSVRG to directly solve the problems with non-smooth component functions, such as SVM. Finally, we empirically study the performance of FSVRG for solving various machine learning problems such as logistic regression, ridge regression, Lasso and SVM. Our results show that FSVRG outperforms the state-of-the-art stochastic methods, including Katyusha.
Corrected a few typos in this version
References in corpus (4)
- SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives
- Stochastic Gradient Descent for Non-smooth Optimization: Convergence Results and Optimal Averaging Schemes
- Accelerated Variance Reduced Stochastic ADMM
- Guaranteed Sufficient Decrease for Variance Reduced Stochastic Gradient Descent
Cited by in corpus (4)
- Quasi-hyperbolic momentum and Adam for deep learning
- Reliability and Performance Assessment of Federated Learning on Clinical Benchmark Data
- Efficient Relaxed Gradient Support Pursuit for Sparsity Constrained Non-convex Optimization
- Larger is Better: The Effect of Learning Rates Enjoyed by Stochastic Optimization with Progressive Variance Reduction