activity
20172021
most citedRethinking gradient sparsification as total error minimization

19 citations · 37 across the 8 of their papers we have counts for

collaborators

12 papers

cs.LG202119 cited

Rethinking gradient sparsification as total error minimization

Atal Narayan Sahu, Aritra Dutta, Ahmed M. Abdelmoniem +3

Gradient compression is a widely-established remedy to tackle the communication bottleneck in distributed training of large deep neural networks (DNNs). Under the error-feedback fr…

cs.LG20216 cited

DeepReduce: A Sparse-tensor Communication Framework for Distributed Deep Learning

Kelly Kostopoulou, Hang Xu, Aritra Dutta +3

Sparse tensors appear frequently in distributed deep learning, either as a direct artifact of the deep neural network's gradients, or as a result of an explicit sparsification proc…

math.OC2020

On the Convergence Analysis of Asynchronous SGD for Solving Consistent Linear Systems

Atal Narayan Sahu, Aritra Dutta, Aashutosh Tiwari +1

In the realm of big data and machine learning, data-parallel, distributed stochastic algorithms have drawn significant attention in the present days.~While the synchronous versions…

cs.DC2019

On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep Learning

Aritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem +4

Compressed communication, in the form of sparsification or quantization of stochastic gradients, is employed to reduce communication costs in distributed data-parallel training of…

math.OC2019

Direct Nonlinear Acceleration

Aritra Dutta, El Houcine Bergou, Yunming Xiao +2

Optimization acceleration techniques such as momentum play a key role in state-of-the-art machine learning algorithms. Recently, generic vector sequence extrapolation techniques, s…

math.OC2019

Best Pair Formulation & Accelerated Scheme for Non-convex Principal Component Pursuit

Aritra Dutta, Filip Hanzely, Jingwei Liang +1

The best pair problem aims to find a pair of points that minimize the distance between two disjoint sets. In this paper, we formulate the classical robust principal component analy…