6 citations · 6 across the 3 of their papers we have counts for
4 papers · 1 filter
Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
Daegun Yoon, Sangyoon Oh
Communication overhead is a major obstacle to scaling distributed training systems. Gradient sparsification is a potential optimization approach to reduce the communication volume…
MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
Daegun Yoon, Sangyoon Oh
Gradient sparsification is a communication optimisation technique for scaling and accelerating distributed deep neural network (DNN) training. It reduces the increasing communicati…
DEFT: Exploiting Gradient Norm Difference between Model Layers for Scalable Gradient Sparsification
Daegun Yoon, Sangyoon Oh
Gradient sparsification is a widely adopted solution for reducing the excessive communication traffic in distributed deep learning. However, most existing gradient sparsifiers have…
Empirical Analysis on Top-k Gradient Sparsification for Distributed Deep Learning in a Supercomputing Environment
Daegun Yoon, Sangyoon Oh
To train deep learning models faster, distributed training on multiple GPUs is the very popular scheme in recent years. However, the communication bandwidth is still a major bottle…