activity
20172025
most citedOn the Convergence of Langevin Monte Carlo: The Interplay between Tail Growth and Smoothness

20 citations · 100 across the 14 of their papers we have counts for

collaborators
Showing stat.MLShow all

19 papers · 1 filter

stat.ML2026

SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws

Konstantinos Christopher Tsiolis, Denny Wu, Christos Thrampoulidis +1

We study the training dynamics of multiclass logistic regression on high-dimensional Gaussian mixture models with a large number of classes and establish precise scaling laws gover…

stat.ML2026

Post-Training with Policy Gradients: Optimality and the Base Model Barrier

Alireza Mousavi-Hosseini, Murat A. Erdogdu

We study post-training linear autoregressive models with outcome and process rewards. Given a context , the model must predict the response

stat.ML2025

Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws

Gérard Ben Arous, Murat A. Erdogdu, Nuri Mert Vural +1

We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where th…

stat.ML2025

When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective

Alireza Mousavi-Hosseini, Clayton Sanford, Denny Wu +1

Theoretical efforts to prove advantages of Transformers in comparison with classical architectures such as feedforward and recurrent neural networks have mostly focused on represen…

stat.ML2024

On the Efficiency of ERM in Feature Learning

Ayoub El Hanchi, Chris J. Maddison, Murat A. Erdogdu

Given a collection of feature maps indexed by a set , we study the performance of empirical risk minimization (ERM) on regression problems with square loss over the un…

stat.ML2024

Robust Feature Learning for Multi-Index Models in High Dimensions

Alireza Mousavi-Hosseini, Adel Javanmard, Murat A. Erdogdu

Recently, there have been numerous studies on feature learning with neural networks, specifically on learning single- and multi-index models where the target is a function of a low…