5 papers · 1 filter
Muon learns balanced solutions in matrix factorization without slow saddle-to-saddle dynamics
Mark Rhee, Jamie Simon, Dhruva Karkada
Matrix factorization (i.e., problems of the form ) is a minimal learning problem that ex…
Predicting kernel regression learning curves from only raw data statistics
Dhruva Karkada, Joseph Turnbull, Yuxi Liu +1
We study kernel regression with common rotation-invariant kernels on real datasets including CIFAR-5m, SVHN, and ImageNet. We give a theoretical framework that predicts learning cu…
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
Daniel Kunin, Giovanni Luca Marchetti, Feng Chen +5
What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dy…
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
Dhruva Karkada, James B. Simon, Yasaman Bahri +1
Self-supervised word embedding algorithms such as word2vec provide a minimal setting for studying representation learning in language modeling. We examine the quartic Taylor approx…
More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
James B. Simon, Dhruva Karkada, Nikhil Ghosh +1
In our era of enormous neural networks, empirical progress has been driven by the philosophy that more is better. Recent deep learning practice has found repeatedly that larger mod…