Accelerated Mini-Batch Stochastic Dual Coordinate Ascent
arXiv:1305.2581
Abstract
Stochastic dual coordinate ascent (SDCA) is an effective technique for solving regularized loss minimization problems in machine learning. This paper considers an extension of SDCA under the mini-batch setting that is often used in practice. Our main contribution is to introduce an accelerated mini-batch version of SDCA and prove a fast convergence rate for this method. We discuss an implementation of our method over a parallel computing system, and compare the results to both the vanilla stochastic dual coordinate ascent and to the accelerated deterministic gradient descent method of \cite{nesterov2007gradient}.
References in corpus (6)
- GraphLab: A New Framework For Parallel Machine Learning
- A Reliable Effective Terascale Linear Learning System
- Parallel Coordinate Descent for L1-Regularized Loss Minimization
- Distributed Coordinate Descent Method for Learning with Big Data
- Better Mini-Batch Algorithms via Accelerated Gradient Methods
- Smooth minimization of nonsmooth functions with parallel coordinate descent methods
Cited by in corpus (16)
- Mini-Batch Semi-Stochastic Gradient Descent in the Proximal Setting
- Distributed Coordinate Descent Method for Learning with Big Data
- An Asynchronous Parallel Stochastic Coordinate Descent Algorithm
- Stochastic Primal-Dual Coordinate Method for Regularized Empirical Risk Minimization
- Adding vs. Averaging in Distributed Primal-Dual Optimization
- Stochastic Dual Ascent for Solving Linear Systems
- Randomized Dual Coordinate Ascent with Arbitrary Sampling
- Smooth minimization of nonsmooth functions with parallel coordinate descent methods
- Primal Method for ERM with Flexible Mini-batching Schemes and Non-convex Losses
- Adaptive Distributed Stochastic Gradient Descent for Minimizing Delay in the Presence of Stragglers
- Sketch and Project: Randomized Iterative Methods for Linear Systems and Inverting Matrices
- Advances in Asynchronous Parallel and Distributed Optimization
- Robust Training in High Dimensions via Block Coordinate Geometric Median Descent
- Improving SAGA via a Probabilistic Interpolation with Gradient Descent
- Learning Under Delayed Feedback: Implicitly Adapting to Gradient Delays
- Stochastic dual averaging methods using variance reduction techniques for regularized empirical risk minimization problems