works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.LG2026

Towards joint scaling laws with optimal batch size schedules

Jiaxiang Li, Zhiqi Bu, Shiyun Xu

The paper derives a theoretical relationship between learning rate and batch size schedules using convex optimization, and proposes a closed‑form optimal batch size schedule that i…

cs.LG2026

Why Muon Outperforms Adam: A Curvature Perspective

Shuche Wang, Fengzhuo Zhang, Jiaxiang Li +2

Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage remains unclear. Our work takes a…

cs.LG2026

Demystifying Manifold Constraints in LLM Pre-training

Kang An, Jiaxiang Li, Donald Goldfarb +1

The empirical success of large language model (LLM) pre-training relies heavily on heuristic stabilization techniques, such as explicit normalization layers and weight decay. While…

math.OC2025

Federated Learning on Riemannian Manifolds: A Gradient-Free Projection-Based Approach

Hongye Wang, Zhaoye Pan, Chang He +2

Federated learning (FL) has emerged as a powerful paradigm for collaborative model training across distributed clients while preserving data privacy. However, existing FL algorithm…

math.OC2025

On Relatively Smooth Optimization over Riemannian Manifolds

Chang He, Jiaxiang Li, Bo Jiang +2

We study optimization over Riemannian embedded submanifolds, where the objective function is relatively smooth in the ambient Euclidean space. Such problems have broad applications…