5 papers
Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?
Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2
AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying…
Nonsmooth Nonconvex-Concave Minimax Optimization: Convergence Criteria and Algorithms
Jinyang Shi, Luo Luo
This paper considers constrained stochastic nonsmooth minimax optimization problem of the form $\min_{\mathbf{x}\in\mathcal{X}}\max_{\mathbf{y}\in\mathcal{Y}}f\left(\mathbf{x},\mat…
Solving Convex-Concave Problems with th-Order Oracle Complexity
Lesi Chen, Xinliang Zhang, Chengchang Liu +3
When the objective has Lipschitz continuous th-order derivatives, it is known that convex-concave minimax problems can be solved with th-order or…
Decentralized Non-convex Stochastic Optimization with Heterogeneous Variance
Hongxu Chen, Ke Wei, Luo Luo
Decentralized optimization is critical for solving large-scale machine learning problems over distributed networks, where multiple nodes collaborate through local communication. In…
Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise
Hongxu Chen, Ke Wei, Xiaoming Yuan +1
The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that…