3 papers
cs.LG2025
Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
Shihong Ding, Haihan Zhang, Hanzhen Zhao +1
In machine learning, the scaling law describes how the model performance improves with the model and data size scaling up. From a learning theory perspective, this class of results…
stat.ML2025
Optimal Algorithms in Linear Regression under Covariate Shift: On the Importance of Precondition
Yuanshi Liu, Haihan Zhang, Qian Chen +1
A common pursuit in modern statistical learning is to attain satisfactory generalization out of the source data distribution (OOD). In theory, the challenge remains unsolved even u…
cs.LG2024
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
Haihan Zhang, Yuanshi Liu, Qianwen Chen +1
Stochastic gradient descent (SGD) is a widely used algorithm in machine learning, particularly for neural network training. Recent studies on SGD for canonical quadratic optimizati…