17 citations · 23 across the 2 of their papers we have counts for
2 papers
cs.LG2021★ 6 cited
The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball Methods
Wei Tao, Sheng Long, Gaowei Wu +1
The adaptive stochastic gradient descent (SGD) with momentum has been widely adopted in deep learning as well as convex optimization. In practice, the last iterate is commonly used…
cs.LG2020★ 17 cited
The Strength of Nesterov's Extrapolation in the Individual Convergence of Nonsmooth Optimization
W. Tao, Z. Pan, G. Wu +1
The extrapolation strategy raised by Nesterov, which can accelerate the convergence rate of gradient descent methods by orders of magnitude when dealing with smooth convex objectiv…