2 papers
cs.LG2025
AdaPM: a Partial Momentum Algorithm for LLM Training
Yimu Zhang, Yuanshi Liu, Cong Fang
In the training of large language models, momentum is widely used and often demonstrated to achieve significant acceleration. However, storing momentum typically presents memory ch…
stat.ML2025
Learning Curves of Stochastic Gradient Descent in Kernel Regression
Haihan Zhang, Weicheng Lin, Yuanshi Liu +1
This paper considers a canonical problem in kernel regression: how good are the model performances when it is trained by the popular online first-order algorithms, compared to the…