6 papers
Two-Layer Linear Auto-Regressive Models Estimate Latent States
Yahya Sattar, Sunmook Choi, Leo Maynard-Zhang +3
Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models learn latent representations remains an op…
Local linear convergence of gradient methods for overparameterized Gaussian mixtures
Jingxing Wang, Vasileios Charisopoulos, Maryam Fazel
We study the problem of learning Gaussian mixture models under overparameterization. Prior work has shown that while overparameterization is essential for avoiding spurious local o…
High-dimensional Limit of SGD for Diagonal Linear Networks
Begoña GarcÃa MalaxechebarrÃa, Courtney Paquette, Maryam Fazel +1
Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet…
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1
We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction.…
Iteratively reweighted kernel machines efficiently learn sparse functions
Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1
The impressive practical performance of neural networks is often attributed to their ability to learn low-dimensional data representations and hierarchical structure directly from…
Spectral norm bound for the product of random Fourier-Walsh matrices
Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1
We consider matrix products of the form , where are normalized random Fourier-Walsh matrices. We identify an interesting polyn…