5 papers
Zeroth-Order Randomized Subspace Newton Methods
Erik Berglund, Sarit Khirirat, Xiaoyu Wang
Zeroth-order methods have become important tools for solving problems where we have access only to function evaluations. However, the zeroth-order methods only using gradient appro…
A flexible framework for communication-efficient machine learning: from HPC to IoT
Sarit Khirirat, Sindri Magnússon, Arda Aytekin +1
With the increasing scale of machine learning tasks, it has become essential to reduce the communication between computing nodes. Early work on gradient compression focused on the…
Compressed Gradient Methods with Hessian-Aided Error Compensation
Sarit Khirirat, Sindri Magnússon, Mikael Johansson
The emergence of big data has caused a dramatic shift in the operating regime for optimization algorithms. The performance bottleneck, which used to be computations, is now often c…
The Convergence of Sparsified Gradient Methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson +3
Distributed training of massive machine learning models, in particular deep neural networks, via Stochastic Gradient Descent (SGD) is becoming commonplace. Several families of comm…
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, Mikael Johansson
Asynchronous computation and gradient compression have emerged as two key techniques for achieving scalability in distributed optimization for large-scale machine learning. This pa…