6 citations · 6 across the 1 of their papers we have counts for
1 paper
Wei Tao, Sheng Long, Gaowei Wu +1
The adaptive stochastic gradient descent (SGD) with momentum has been widely adopted in deep learning as well as convex optimization. In practice, the last iterate is commonly used…