108 citations · 128 across the 4 of their papers we have counts for
Showing 2018Show all
2 papers · 1 filter
cs.LG2018★ 108 cited
Gradient Descent Happens in a Tiny Subspace
Guy Gur-Ari, Daniel A. Roberts, Ethan Dyer
We show that in a variety of large-scale deep learning scenarios the gradient dynamically converges to a very small subspace after a short period of training. The subspace is spann…
hep-th2018
Operator growth in the SYK model
Daniel A. Roberts, Douglas Stanford, Alexandre Streicher
We discuss the probability distribution for the "size" of a time-evolving operator in the SYK model. Scrambling is related to the fact that as time passes, the distribution shifts…