Partitioning Data on Features or Samples in Communication-Efficient Distributed Optimization?
arXiv:1510.06688
Abstract
In this paper we study the effect of the way that the data is partitioned in distributed optimization. The original DiSCO algorithm [Communication-Efficient Distributed Optimization of Self-Concordant Empirical Loss, Yuchen Zhang and Lin Xiao, 2015] partitions the input data based on samples. We describe how the original algorithm has to be modified to allow partitioning on features and show its efficiency both in theory and also in practice.
References in corpus (5)
- Communication Efficient Distributed Optimization using an Approximate Newton-type Method
- Parallel Coordinate Descent for L1-Regularized Loss Minimization
- Distributed Block Coordinate Descent for Minimizing Partially Separable Functions
- Distributed Mini-Batch SDCA
- Analysis of Distributed Stochastic Dual Coordinate Ascent
Cited by in corpus (5)
- Distributed Learning with Compressed Gradient Differences
- Stochastic, Distributed and Federated Optimization for Machine Learning
- Communication trade-offs for synchronized distributed SGD with large step size
- Communication Optimality Trade-offs For Distributed Estimation
- Clustering with Distributed Data