collaborators

5 papers

cs.LG2026

Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2

AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying…

math.OC2026

Nonsmooth Nonconvex-Concave Minimax Optimization: Convergence Criteria and Algorithms

Jinyang Shi, Luo Luo

This paper considers constrained stochastic nonsmooth minimax optimization problem of the form $\min_{\mathbf{x}\in\mathcal{X}}\max_{\mathbf{y}\in\mathcal{Y}}f\left(\mathbf{x},\mat…

math.OC2026

Solving Convex-Concave Problems with th-Order Oracle Complexity

Lesi Chen, Xinliang Zhang, Chengchang Liu +3

When the objective has Lipschitz continuous th-order derivatives, it is known that convex-concave minimax problems can be solved with th-order or…

math.OC2026

Decentralized Non-convex Stochastic Optimization with Heterogeneous Variance

Hongxu Chen, Ke Wei, Luo Luo

Decentralized optimization is critical for solving large-scale machine learning problems over distributed networks, where multiple nodes collaborate through local communication. In…

cs.LG2026

Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise

Hongxu Chen, Ke Wei, Xiaoming Yuan +1

The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that…