17 citations · 23 across the 4 of their papers we have counts for
4 papers
A Convergence Analysis of Nesterov's Accelerated Gradient Method in Training Deep Linear Neural Networks
Xin Liu, Wei Tao, Zhisong Pan
Momentum methods, including heavy-ball~(HB) and Nesterov's accelerated gradient~(NAG), are widely used in training neural networks for their fast convergence. However, there is a l…
The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball Methods
Wei Tao, Sheng Long, Gaowei Wu +1
The adaptive stochastic gradient descent (SGD) with momentum has been widely adopted in deep learning as well as convex optimization. In practice, the last iterate is commonly used…
Gradient Descent Averaging and Primal-dual Averaging for Strongly Convex Optimization
Wei Tao, Wei Li, Zhisong Pan +1
Averaging scheme has attracted extensive attention in deep learning as well as traditional machine learning. It achieves theoretically optimal convergence and also improves the emp…
The Strength of Nesterov's Extrapolation in the Individual Convergence of Nonsmooth Optimization
W. Tao, Z. Pan, G. Wu +1
The extrapolation strategy raised by Nesterov, which can accelerate the convergence rate of gradient descent methods by orders of magnitude when dealing with smooth convex objectiv…