8 citations · 8 across the 1 of their papers we have counts for
1 paper
Xiaozhe Gu, Zixun Zhang, Yuncheng Jiang +4
Despite the simplicity, stochastic gradient descent (SGD)-like algorithms are successful in training deep neural networks (DNNs). Among various attempts to improve SGD, weight aver…