Smooth minimization of nonsmooth functions with parallel coordinate descent methods
arXiv:1309.5885 · doi:10.1007/978-3-030-12119-8_4
Abstract
We study the performance of a family of randomized parallel coordinate descent methods for minimizing the sum of a nonsmooth and separable convex functions. The problem class includes as a special case L1-regularized L1 regression and the minimization of the exponential loss ("AdaBoost problem"). We assume the input data defining the loss function is contained in a sparse matrix with at most nonzeros in each row. Our methods need iterations to find an approximate solution with high probability, where is the number of processors and for the fastest variant. The notation hides dependence on quantities such as the required accuracy and confidence levels and the distance of the starting iterate from an optimal point. Since is a decreasing function of , the method needs fewer iterations when more processors are used. Certain variants of our algorithms perform on average only $O(\nnz(A)/n)$ arithmetic operations during a single iteration per processor and, because decreases when does, fewer iterations are needed for sparser problems.
39 pages, 1 algorithm, 3 figures, 2 tables
References in corpus (9)
- Generalized power method for sparse principal component analysis
- Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization
- Parallel Coordinate Descent for L1-Regularized Loss Minimization
- Block-Coordinate Frank-Wolfe Optimization for Structural SVMs
- Parallel Coordinate Descent Methods for Big Data Optimization
- Accelerated Mini-Batch Stochastic Dual Coordinate Ascent
- Iteration Complexity of Randomized Block-Coordinate Descent Methods for Minimizing a Composite Function
- Parallel coordinate descent for the Adaboost problem
- Stochastic Block Mirror Descent Methods for Nonsmooth and Stochastic Optimization
Cited by in corpus (24)
- Distributed Coordinate Descent Method for Learning with Big Data
- Parallel Coordinate Descent Methods for Big Data Optimization
- Accelerated Mini-Batch Stochastic Dual Coordinate Ascent
- Parallel Successive Convex Approximation for Nonsmooth Nonconvex Optimization
- A Primer on Coordinate Descent Algorithms
- Stochastic Dual Ascent for Solving Linear Systems
- Distributed Block Coordinate Descent for Minimizing Partially Separable Functions
- Randomized Dual Coordinate Ascent with Arbitrary Sampling
- SDNA: Stochastic Dual Newton Ascent for Empirical Risk Minimization
- Accelerated, Parallel and Proximal Coordinate Descent
- Coordinate Descent with Arbitrary Sampling I: Algorithms and Complexity
- Stochastic, Distributed and Federated Optimization for Machine Learning
- Stochastic Dual Coordinate Ascent with Adaptive Probabilities
- Restarting accelerated gradient methods with a rough strong convexity estimate
- On Optimal Probabilities in Stochastic Coordinate Descent Methods
- Coordinate Descent with Arbitrary Sampling II: Expected Separable Overapproximation
- Privacy Preserving Randomized Gossip Algorithms
- Parallel coordinate descent for the Adaboost problem
- Asynchronous Stochastic Coordinate Descent: Parallelism and Convergence Properties
- Stochastic Coordinate Minimization with Progressive Precision for Stochastic Convex Optimization
- Smooth Primal-Dual Coordinate Descent Algorithms for Nonsmooth Convex Optimization
- Data Sampling Strategies in Stochastic Algorithms for Empirical Risk Minimization
- Efficient random coordinate descent algorithms for large-scale structured nonconvex optimization
- A generic coordinate descent solver for nonsmooth convex optimization