5 papers
Adam Converges in Nonsmooth Nonconvex Optimization
Zijian Liu
Adam is one of the most widely implemented and influential modern optimizers. Why is it effective across different optimization problems in practice? This question arguably lies at…
In-Expectation Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise
Zijian Liu
Many stochastic gradient methods are believed not to converge when the noise in stochastic gradients has only a finite -th moment for , a setting known as…
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
Zijian Liu
Many tasks in modern machine learning are observed to involve heavy-tailed gradient noise during the optimization process. To manage this realistic and challenging setting, new mec…
Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis
Zijian Liu
Optimization under heavy-tailed noise has become popular recently, since it better fits many modern machine learning tasks, as captured by empirical observations. Concretely, inste…
Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
Zijian Liu
In Online Convex Optimization (OCO), when the stochastic gradient has a finite variance, many algorithms provably work and guarantee a sublinear regret. However, limited results ar…