Mini-Batch Primal and Dual Methods for SVMs
arXiv:1303.2314
Abstract
We address the issue of using mini-batches in stochastic optimization of SVMs. We show that the same quantity, the spectral norm of the data, controls the parallelization speedup obtained for both primal stochastic subgradient descent (SGD) and stochastic dual coordinate ascent (SCDA) methods and use it to derive novel variants of mini-batched SDCA. Our guarantees for both methods are expressed in terms of the original nonsmooth primal problem based on the hinge-loss.
References in corpus (1)
Cited by in corpus (6)
- Federated Optimization: Distributed Machine Learning for On-Device Intelligence
- Stochastic Dual Coordinate Ascent with Adaptive Probabilities
- Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
- Coordinate Descent with Arbitrary Sampling II: Expected Separable Overapproximation
- Block-diagonal Hessian-free Optimization for Training Neural Networks
- Data Dependent Convergence for Distributed Stochastic Optimization