108 citations · 128 across the 4 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2022★ 12 cited
A Solvable Model of Neural Scaling Laws
Alexander Maloney, Daniel A. Roberts, James Sully
Large language models with a huge number of parameters, when trained on near internet-sized number of tokens, have been empirically shown to obey neural scaling laws: specifically,…
cs.LG2021★ 7 cited
SGD Implicitly Regularizes Generalization Error
Daniel A. Roberts
We derive a simple and model-independent formula for the change in the generalization gap due to a gradient descent update. We then compare the change in the test error for stochas…
cs.LG2018★ 108 cited
Gradient Descent Happens in a Tiny Subspace
Guy Gur-Ari, Daniel A. Roberts, Ethan Dyer
We show that in a variety of large-scale deep learning scenarios the gradient dynamically converges to a very small subspace after a short period of training. The subspace is spann…