Showing stat.MLShow all
2 papers · 1 filter
stat.ML2024
No Free Prune: Information-Theoretic Barriers to Pruning at Initialization
Tanishq Kumar, Kevin Luo, Mark Sellke
The existence of "lottery tickets" arXiv:1803.03635 at or near initialization raises the tantalizing question of whether large models are necessary in deep learning, or whether spa…
stat.ML2024
Grokking as the Transition from Lazy to Rich Training Dynamics
Tanishq Kumar, Blake Bordelon, Samuel J. Gershman +1
We propose that the grokking phenomenon, where the train loss of a neural network decreases much earlier than its test loss, can arise due to a neural network transitioning from la…