2 papers
cs.LG2026
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
Diyuan Wu, Lehan Chen, Theodor Misiakiewicz +1
It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenomenon of weak-to-strong generaliz…
math.NA2024
Eigen-componentwise convergence of SGD on quadratic programming
Lehan Chen, Yuji Nakatsukasa
Stochastic gradient descent (SGD) is a workhorse algorithm for solving large-scale optimization problems in data science and machine learning. Understanding the convergence of SGD…