6 papers
Stochastic Saddle Avoidance Beyond Unit Excitation and Smoothness: A Pathwise Lyapunov-Perron Framework
Junwen Qiu, Bohao Ma, Andre Milzarek +1
Unit excitation (UE) is a common assumption in stochastic saddle avoidance: the stochastic error must have a uniformly positive component along every direction, in expectation. Thi…
Random Reshuffling with Momentum: Complexity Bounds and Last-iterate Convergence
Junwen Qiu, Bohao Ma, Andre Milzarek
Random reshuffling with momentum (RRM) corresponds to the SGD optimizer with the 'momentum' option enabled, as found in many machine learning libraries such as PyTorch and TensorFl…
A Normal Map-Based Proximal Stochastic Gradient Method: Convergence and Identification Properties
Junwen Qiu, Li Jiang, Andre Milzarek
The proximal stochastic gradient method (PSGD) is one of the state-of-the-art approaches for stochastic composite-type problems. In contrast to its deterministic counterpart, PSGD…
A Generalized Version of Chung's Lemma and its Applications
Li Jiang, Xiao Li, Andre Milzarek +1
Chung's Lemma is a classical tool for establishing asymptotic convergence rates of (stochastic) optimization methods under strong convexity-type assumptions and appropriate polynom…
A New Random Reshuffling Method for Nonsmooth Nonconvex Finite-sum Optimization
Junwen Qiu, Xiao Li, Andre Milzarek
Random reshuffling techniques are prevalent in large-scale applications, such as training neural networks. While the convergence and acceleration effects of random reshuffling-type…
Convergence of SGD with momentum in the nonconvex case: A time window-based analysis
Junwen Qiu, Bohao Ma, Andre Milzarek
The stochastic gradient descent method with momentum (SGDM) is a common approach for solving large-scale and stochastic optimization problems. Despite its popularity, the convergen…