Parallel Coordinate Descent for L1-Regularized Loss Minimization
arXiv:1105.5379
Abstract
We propose Shotgun, a parallel coordinate descent algorithm for minimizing L1-regularized losses. Though coordinate descent seems inherently sequential, we prove convergence bounds for Shotgun which predict linear speedups, up to a problem-dependent limit. We present a comprehensive empirical study of Shotgun for Lasso and sparse logistic regression. Our theoretical predictions on the potential for parallelism closely match behavior on real data. Shotgun outperforms other published solvers on a range of large problems, proving to be one of the most scalable algorithms for L1.
Cited by in corpus (20)
- Geotagging One Hundred Million Twitter Accounts with Total Variation Minimization
- Communication-Efficient Distributed Dual Coordinate Ascent
- Mini-Batch Primal and Dual Methods for SVMs
- Parallel Successive Convex Approximation for Nonsmooth Nonconvex Optimization
- Randomized Dual Coordinate Ascent with Arbitrary Sampling
- PASSCoDe: Parallel ASynchronous Stochastic dual Co-ordinate Descent
- Feature Clustering for Accelerating Parallel Coordinate Descent
- An Accelerated Proximal Coordinate Gradient Method and its Application to Regularized Empirical Risk Minimization
- Randomized Block Coordinate Descent for Online and Stochastic Optimization
- Scaling Up Coordinate Descent Algorithms for Large Regularization Problems
- Median Selection Subset Aggregation for Parallel Inference
- LOCO: Distributing Ridge Regression with Random Projections
- Coordinate Descent with Arbitrary Sampling II: Expected Separable Overapproximation
- Primitives for Dynamic Big Model Parallelism
- Large-scale randomized-coordinate descent methods with non-separable linear constraints
- A machine-compiled macroevolutionary history of Phanerozoic life
- Accelerated Parallel Optimization Methods for Large Scale Machine Learning
- A Reduction of the Elastic Net to Support Vector Machines with an Application to GPU Computing
- A Parallel and Efficient Algorithm for Learning to Match
- Distributed Optimization via Adaptive Regularization for Large Problems with Separable Constraints