5 papers
Preconditioning Benefits of Spectral Orthogonalization in Muon
Jianhao Ma, Yu Huang, Yuejie Chi +1
The Muon optimizer, a matrix-structured algorithm that leverages spectral orthogonalization of gradients, is a milestone in the pretraining of large language models. However, the u…
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
Yu Huang, Zixin Wen, Aarti Singh +2
The ability to reason lies at the core of artificial intelligence (AI), and challenging problems usually call for deeper and longer reasoning to tackle. A crucial question about AI…
Statistical and Algorithmic Foundations of Reinforcement Learning
Yuejie Chi, Yuxin Chen, Yuting Wei
As a paradigm for sequential decision making in unknown environments, reinforcement learning (RL) has received a flurry of attention in recent years. However, the explosion of mode…
Transformers Meet In-Context Learning: A Universal Approximation Theory
Gen Li, Yuchen Jiao, Yu Huang +2
Large language models are capable of in-context learning, the ability to perform new tasks at test time using a handful of input-output examples, without parameter updates. We deve…
Stochastic Runge-Kutta Methods: Provable Acceleration of Diffusion Models
Yuchen Wu, Yuxin Chen, Yuting Wei
Diffusion models play a pivotal role in contemporary generative modeling, claiming state-of-the-art performance across various domains. Despite their superior sample quality, mains…