5 papers
High Probability Guarantees for Random Reshuffling
Hengxu Yu, Xiao Li
We consider the stochastic gradient method with random reshuffling () for tackling smooth nonconvex optimization problems. finds broad applications in pr…
A Generalized Version of Chung's Lemma and its Applications
Li Jiang, Xiao Li, Andre Milzarek +1
Chung's Lemma is a classical tool for establishing asymptotic convergence rates of (stochastic) optimization methods under strong convexity-type assumptions and appropriate polynom…
A New Random Reshuffling Method for Nonsmooth Nonconvex Finite-sum Optimization
Junwen Qiu, Xiao Li, Andre Milzarek
Random reshuffling techniques are prevalent in large-scale applications, such as training neural networks. While the convergence and acceleration effects of random reshuffling-type…
Dual Acceleration for Minimax Optimization: Linear Convergence Under Relaxed Assumptions
Jingwang Li, Xiao Li
This paper addresses the bilinearly coupled minimax optimization problem: $\min_{x \in \mathbb{R}^{d_x}}\max_{y \in \mathbb{R}^{d_y}} \ f_1(x) + f_2(x) + y^{\top} Bx - g_1(y) - g_2…
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
Ruinan Jin, Xiao Li, Yaoliang Yu +1
Adaptive Moment Estimation (Adam) is a cornerstone optimization algorithm in deep learning, widely recognized for its flexibility with adaptive learning rates and efficiency in han…