341 citations · 717 across the 109 of their papers we have counts for
8 papers · 2 filters
Byzantine-Resilient Non-Convex Stochastic Gradient Descent
Zeyuan Allen-Zhu, Faeze Ebrahimian, Jerry Li +1
We study adversary-resilient stochastic distributed optimization, in which machines can independently compute stochastic gradients, and cooperate to jointly optimize over their…
Adaptive Gradient Quantization for Data-Parallel SGD
Fartash Faghri, Iman Tabrizian, Ilia Markov +3
Many communication-efficient variants of SGD use gradient quantization schemes. These schemes are often heuristic and fixed over the course of training. We empirically observe that…
Towards Tight Communication Lower Bounds for Distributed Optimisation
Dan Alistarh, Janne H. Korhonen
We consider a standard distributed optimisation setting where machines, each holding a -dimensional function , aim to jointly minimise the sum of the functions $\sum_{i…
Stochastic Gradient Langevin with Delayed Gradients
Vyacheslav Kungurtsev, Bapi Chatterjee, Dan Alistarh
Stochastic Gradient Langevin Dynamics (SGLD) ensures strong guarantees with regards to convergence in measure for sampling log-concave posterior distributions by adding noise to st…
WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
Sidak Pal Singh, Dan Alistarh
Second-order information, in the form of Hessian- or Inverse-Hessian-vector products, is a fundamental tool for solving optimization problems. Recently, there has been significant…
On the Sample Complexity of Adversarial Multi-Source PAC Learning
Nikola Konstantinov, Elias Frantar, Dan Alistarh +1
We study the problem of learning from multiple untrusted data sources, a scenario of increasing practical relevance given the recent emergence of crowdsourcing and collaborative le…