activity
20242026
collaborators

10 papers

math.OC2026

A lower bound for stepsize-based acceleration of gradient descent

Jianhao Ma, Yuxin Chen

Recent work has shown that, for smooth convex optimization, plain gradient descent can be accelerated from its textbook convergence rate of (where denotes the numbe…

cs.LG2026

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions

Yuepeng Yang, Yuxin Chen, Yuejie Chi

Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and…

cs.LG2026

WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training

Jianhao Ma, Yuxin Chen

Standard learning rate schedules such as cosine annealing are tied to a fixed training horizon, limiting their ability to accommodate post hoc horizon extension. Warmup-stable-deca…

cs.LG2026

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

Yu Huang, Zixin Wen, Yuejie Chi +4

Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on…

cs.LG2026

Denoising diffusion probabilistic models are optimally adaptive to unknown low dimensionality

Zhihan Huang, Yuting Wei, Yuxin Chen

The denoising diffusion probabilistic model (DDPM) has emerged as a mainstream generative model in generative AI. While sharp convergence guarantees have been established for the D…

cs.LG2026

Preconditioning Benefits of Spectral Orthogonalization in Muon

Jianhao Ma, Yu Huang, Yuejie Chi +1

The Muon optimizer, a matrix-structured algorithm that leverages spectral orthogonalization of gradients, is a milestone in the pretraining of large language models. However, the u…