20 citations · 100 across the 14 of their papers we have counts for
19 papers · 1 filter
SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws
Konstantinos Christopher Tsiolis, Denny Wu, Christos Thrampoulidis +1
We study the training dynamics of multiclass logistic regression on high-dimensional Gaussian mixture models with a large number of classes and establish precise scaling laws gover…
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
Alireza Mousavi-Hosseini, Murat A. Erdogdu
We study post-training linear autoregressive models with outcome and process rewards. Given a context , the model must predict the response …
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
Gérard Ben Arous, Murat A. Erdogdu, Nuri Mert Vural +1
We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where th…
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
Alireza Mousavi-Hosseini, Clayton Sanford, Denny Wu +1
Theoretical efforts to prove advantages of Transformers in comparison with classical architectures such as feedforward and recurrent neural networks have mostly focused on represen…
On the Efficiency of ERM in Feature Learning
Ayoub El Hanchi, Chris J. Maddison, Murat A. Erdogdu
Given a collection of feature maps indexed by a set , we study the performance of empirical risk minimization (ERM) on regression problems with square loss over the un…
Robust Feature Learning for Multi-Index Models in High Dimensions
Alireza Mousavi-Hosseini, Adel Javanmard, Murat A. Erdogdu
Recently, there have been numerous studies on feature learning with neural networks, specifically on learning single- and multi-index models where the target is a function of a low…