From the 1 of 5 linked papers with an AI index.
5 papers
Towards joint scaling laws with optimal batch size schedules
Jiaxiang Li, Zhiqi Bu, Shiyun Xu
The paper derives a theoretical relationship between learning rate and batch size schedules using convex optimization, and proposes a closed‑form optimal batch size schedule that i…
Why Muon Outperforms Adam: A Curvature Perspective
Shuche Wang, Fengzhuo Zhang, Jiaxiang Li +2
Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage remains unclear. Our work takes a…
Demystifying Manifold Constraints in LLM Pre-training
Kang An, Jiaxiang Li, Donald Goldfarb +1
The empirical success of large language model (LLM) pre-training relies heavily on heuristic stabilization techniques, such as explicit normalization layers and weight decay. While…
Federated Learning on Riemannian Manifolds: A Gradient-Free Projection-Based Approach
Hongye Wang, Zhaoye Pan, Chang He +2
Federated learning (FL) has emerged as a powerful paradigm for collaborative model training across distributed clients while preserving data privacy. However, existing FL algorithm…
On Relatively Smooth Optimization over Riemannian Manifolds
Chang He, Jiaxiang Li, Bo Jiang +2
We study optimization over Riemannian embedded submanifolds, where the objective function is relatively smooth in the ambient Euclidean space. Such problems have broad applications…