collaborators

6 papers

cs.LG2026

Two-Layer Linear Auto-Regressive Models Estimate Latent States

Yahya Sattar, Sunmook Choi, Leo Maynard-Zhang +3

Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models learn latent representations remains an op…

cs.LG2026

Local linear convergence of gradient methods for overparameterized Gaussian mixtures

Jingxing Wang, Vasileios Charisopoulos, Maryam Fazel

We study the problem of learning Gaussian mixture models under overparameterization. Prior work has shown that while overparameterization is essential for avoiding spurious local o…

math.OC2026

High-dimensional Limit of SGD for Diagonal Linear Networks

Begoña García Malaxechebarría, Courtney Paquette, Maryam Fazel +1

Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet…

stat.ML2026

Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models

Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1

We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction.…

stat.ML2025

Iteratively reweighted kernel machines efficiently learn sparse functions

Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1

The impressive practical performance of neural networks is often attributed to their ability to learn low-dimensional data representations and hierarchical structure directly from…

math.PR2025

Spectral norm bound for the product of random Fourier-Walsh matrices

Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1

We consider matrix products of the form , where are normalized random Fourier-Walsh matrices. We identify an interesting polyn…