4 papers
Near-Optimal Tensor PCA via Normalized Stochastic Gradient Ascent with Overparameterization
Shihong Ding, Yihong Gu, Yuanshi Liu +1
We study the Order- () spiked tensor model for the tensor principal component analysis (PCA) problem: given i.i.d. observations of a -th order tensor generated…
AdaPM: a Partial Momentum Algorithm for LLM Training
Yimu Zhang, Yuanshi Liu, Cong Fang
In the training of large language models, momentum is widely used and often demonstrated to achieve significant acceleration. However, storing momentum typically presents memory ch…
Learning Curves of Stochastic Gradient Descent in Kernel Regression
Haihan Zhang, Weicheng Lin, Yuanshi Liu +1
This paper considers a canonical problem in kernel regression: how good are the model performances when it is trained by the popular online first-order algorithms, compared to the…
Optimal Algorithms in Linear Regression under Covariate Shift: On the Importance of Precondition
Yuanshi Liu, Haihan Zhang, Qian Chen +1
A common pursuit in modern statistical learning is to attain satisfactory generalization out of the source data distribution (OOD). In theory, the challenge remains unsolved even u…