1 citations · 1 across the 3 of their papers we have counts for
8 papers
Muon learns balanced solutions in matrix factorization without slow saddle-to-saddle dynamics
Mark Rhee, Jamie Simon, Dhruva Karkada
Matrix factorization (i.e., problems of the form ) is a minimal learning problem that ex…
Symmetry in language statistics shapes the geometry of model representations
Dhruva Karkada, Daniel J. Korchinski, Andres Nava +2
The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth on…
There Will Be a Scientific Theory of Deep Learning
Jamie Simon, Daniel Kunin, Alexander Atanasov +11
In this paper, we make the case that a scientific theory of deep learning is emerging. By this we mean a theory which characterizes important properties and statistics of the train…
Predicting kernel regression learning curves from only raw data statistics
Dhruva Karkada, Joseph Turnbull, Yuxi Liu +1
We study kernel regression with common rotation-invariant kernels on real datasets including CIFAR-5m, SVHN, and ImageNet. We give a theoretical framework that predicts learning cu…
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
Daniel Kunin, Giovanni Luca Marchetti, Feng Chen +5
What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dy…
On the Emergence of Linear Analogies in Word Embeddings
Daniel J. Korchinski, Dhruva Karkada, Yasaman Bahri +1
Models such as Word2Vec and GloVe construct word embeddings based on the co-occurrence probability of words and in text corpora. The resulting vectors not on…