4 papers
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
Behrad Moniri, Hamed Hassani
We study feature learning in two-layer neural networks within the linear-width regime, where the number of hidden neurons, sample size, and input dimension scale proportionally. Wh…
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
Behrad Moniri, Hamed Hassani
Weak-to-strong generalization, where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher, has been widely observed but the…
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
Thomas T. Zhang, Behrad Moniri, Ansh Nagwekar +4
Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods ha…
Asymptotics of Linear Regression with Linearly Dependent Data
Behrad Moniri, Hamed Hassani
In this paper we study the asymptotics of linear regression in settings with non-Gaussian covariates where the covariates exhibit a linear dependency structure, departing from the…