3 citations · 5 across the 12 of their papers we have counts for
3 papers · 1 filter
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
Neil Mallinar, Daniel Beaglehole, Libin Zhu +3
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accura…
Linear Recursive Feature Machines provably recover low-rank matrices
Adityanarayanan Radhakrishnan, Mikhail Belkin, Dmitriy Drusvyatskiy
A fundamental problem in machine learning is to understand how neural networks make accurate predictions, while seemingly bypassing the curse of dimensionality. A possible explanat…
Mechanism of feature learning in convolutional neural networks
Daniel Beaglehole, Adityanarayanan Radhakrishnan, Parthe Pandit +1
Understanding the mechanism of how convolutional neural networks learn features from image data is a fundamental problem in machine learning and computer vision. In this work, we i…