108 citations · 116 across the 3 of their papers we have counts for
5 papers
SGD Implicitly Regularizes Generalization Error
Daniel A. Roberts
We derive a simple and model-independent formula for the change in the generalization gap due to a gradient descent update. We then compare the change in the test error for stochas…
Why is AI hard and Physics simple?
Daniel A. Roberts
We discuss why AI is hard and why physics is simple. We discuss how physical intuition and the approach of theoretical physics can be brought to bear on the field of artificial int…
Robust Learning with Jacobian Regularization
Judy Hoffman, Daniel A. Roberts, Sho Yaida
Design of reliable systems must guarantee stability against input perturbations. In machine learning, such guarantee entails preventing overfitting and ensuring robustness of model…
Gradient Descent Happens in a Tiny Subspace
Guy Gur-Ari, Daniel A. Roberts, Ethan Dyer
We show that in a variety of large-scale deep learning scenarios the gradient dynamically converges to a very small subspace after a short period of training. The subspace is spann…
Operator growth in the SYK model
Daniel A. Roberts, Douglas Stanford, Alexandre Streicher
We discuss the probability distribution for the "size" of a time-evolving operator in the SYK model. Scrambling is related to the fact that as time passes, the distribution shifts…