2 papers
cs.LG2018
SmoothOut: Smoothing Out Sharp Minima to Improve Generalization in Deep Learning
Wei Wen, Yandan Wang, Feng Yan +4
In Deep Learning, Stochastic Gradient Descent (SGD) is usually selected as a training method because of its efficiency; however, recently, a problem in SGD gains research interest:…
cs.LG2017
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
Wei Wen, Cong Xu, Feng Yan +4
High network communication cost for synchronizing gradients and parameters is the well-known bottleneck of distributed training. In this work, we propose TernGrad that uses ternary…