3 papers
stat.ML2026
Asymmetric Scaling Laws from Sparse Features
John Sous, Michael Winer
We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input.…
cs.LG2026
Estimating the expected output of wide random MLPs more efficiently than sampling
Wilson Wu, Victor Lecomte, Michael Winer +3
By far the most common way to estimate an expected loss in machine learning is to draw samples, compute the loss on each one, and take the empirical average. However, sampling is n…
cond-mat.dis-nn2025
Deep Neural Nets as Hamiltonians
Mike Winer, Boris Hanin
Neural networks are complex functions of both their inputs and parameters. Much prior work in deep learning theory analyzes the distribution of network outputs at a fixed a set of…