6 papers
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
Behrad Moniri, Hamed Hassani
We study feature learning in two-layer neural networks within the linear-width regime, where the number of hidden neurons, sample size, and input dimension scale proportionally. Wh…
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
Behrad Moniri, Hamed Hassani
Weak-to-strong generalization, where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher, has been widely observed but the…
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
Behrad Moniri, Donghwan Lee, Hamed Hassani +1
Feature learning is thought to be one of the fundamental reasons for the success of deep neural networks. It is rigorously known that in two-layer fully-connected neural networks u…
Evaluating the Performance of Large Language Models via Debates
Behrad Moniri, Hamed Hassani, Edgar Dobriban
Large Language Models (LLMs) are rapidly evolving and impacting various fields, necessitating the development of effective methods to evaluate and compare their performance. Most c…
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
Thomas T. Zhang, Behrad Moniri, Ansh Nagwekar +4
Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods ha…
Asymptotics of Linear Regression with Linearly Dependent Data
Behrad Moniri, Hamed Hassani
In this paper we study the asymptotics of linear regression in settings with non-Gaussian covariates where the covariates exhibit a linear dependency structure, departing from the…