A Parallel SGD method with Strong Convergence
arXiv:1311.0636
Abstract
This paper proposes a novel parallel stochastic gradient descent (SGD) method that is obtained by applying parallel sets of SGD iterations (each set operating on one node using the data residing in it) for finding the direction in each iteration of a batch descent method. The method has strong convergence properties. Experiments on datasets with high dimensional feature spaces show the value of this method.
Cited by in corpus (4)
- Communication Efficient Distributed Optimization using an Approximate Newton-type Method
- Communication Complexity of Distributed Convex Learning and Optimization
- GIANT: Globally Improved Approximate Newton Method for Distributed Optimization
- Without-Replacement Sampling for Stochastic Gradient Methods: Convergence Results and Application to Distributed Optimization