10 papers
A lower bound for stepsize-based acceleration of gradient descent
Jianhao Ma, Yuxin Chen
Recent work has shown that, for smooth convex optimization, plain gradient descent can be accelerated from its textbook convergence rate of (where denotes the numbe…
Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions
Yuepeng Yang, Yuxin Chen, Yuejie Chi
Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and…
WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training
Jianhao Ma, Yuxin Chen
Standard learning rate schedules such as cosine annealing are tied to a fixed training horizon, limiting their ability to accommodate post hoc horizon extension. Warmup-stable-deca…
On the Emergence of Implicit Curriculum in RLVR Learning Dynamics
Yu Huang, Zixin Wen, Yuejie Chi +4
Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on…
Denoising diffusion probabilistic models are optimally adaptive to unknown low dimensionality
Zhihan Huang, Yuting Wei, Yuxin Chen
The denoising diffusion probabilistic model (DDPM) has emerged as a mainstream generative model in generative AI. While sharp convergence guarantees have been established for the D…
Preconditioning Benefits of Spectral Orthogonalization in Muon
Jianhao Ma, Yu Huang, Yuejie Chi +1
The Muon optimizer, a matrix-structured algorithm that leverages spectral orthogonalization of gradients, is a milestone in the pretraining of large language models. However, the u…