2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2021
Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent
Kun Zeng, Jinlan Liu, Zhixia Jiang +1
The plain stochastic gradient descent and momentum stochastic gradient descent have extremely wide applications in deep learning due to their simple settings and low computational…
cs.LG2021★ 2 cited
A decreasing scaling transition scheme from Adam to SGD
Kun Zeng, Jinlan Liu, Zhixia Jiang +1
Adaptive gradient algorithm (AdaGrad) and its variants, such as RMSProp, Adam, AMSGrad, etc, have been widely used in deep learning. Although these algorithms are faster in the ear…