activity
20242026
collaborators

6 papers

stat.ML2026

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent

Behrad Moniri, Hamed Hassani

We study feature learning in two-layer neural networks within the linear-width regime, where the number of hidden neurons, sample size, and input dimension scale proportionally. Wh…

stat.ML2025

On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective

Behrad Moniri, Hamed Hassani

Weak-to-strong generalization, where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher, has been widely observed but the…

stat.ML2025

A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks

Behrad Moniri, Donghwan Lee, Hamed Hassani +1

Feature learning is thought to be one of the fundamental reasons for the success of deep neural networks. It is rigorously known that in two-layer fully-connected neural networks u…

cs.CL2025

Evaluating the Performance of Large Language Models via Debates

Behrad Moniri, Hamed Hassani, Edgar Dobriban

Large Language Models (LLMs) are rapidly evolving and impacting various fields, necessitating the development of effective methods to evaluate and compare their performance. Most c…

cs.LG2025

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

Thomas T. Zhang, Behrad Moniri, Ansh Nagwekar +4

Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods ha…

stat.ML2024

Asymptotics of Linear Regression with Linearly Dependent Data

Behrad Moniri, Hamed Hassani

In this paper we study the asymptotics of linear regression in settings with non-Gaussian covariates where the covariates exhibit a linear dependency structure, departing from the…