activity
20142020
most citedGradient Descent Happens in a Tiny Subspace

108 citations · 269 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG202014 cited

Asymptotics of Wide Convolutional Neural Networks

Anders Andreassen, Ethan Dyer

Wide neural networks have proven to be a rich class of architectures for both theory and practice. Motivated by the observation that finite width convolutional networks appear to o…

stat.ML202058 cited

The large learning rate phase of deep learning: the catapult mechanism

Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer +2

The choice of initial learning rate can have a profound effect on the performance of deep networks. We present a class of neural networks with solvable training dynamics, and confi…

cs.LG201945 cited

Asymptotics of Wide Networks from Feynman Diagrams

Ethan Dyer, Guy Gur-Ari

Understanding the asymptotic behavior of wide networks is of considerable interest. In this work, we present a general method for analyzing this large width behavior. The method is…

cs.LG2018108 cited

Gradient Descent Happens in a Tiny Subspace

Guy Gur-Ari, Daniel A. Roberts, Ethan Dyer

We show that in a variety of large-scale deep learning scenarios the gradient dynamically converges to a very small subspace after a short period of training. The subspace is spann…

hep-th201444 cited

Super-Rényi Entropy & Wilson Loops for N=4 SYM and their Gravity Duals

Michael Crossley, Ethan Dyer, Julian Sonner

We compute the supersymmetric Rényi entropies across a spherical entanglement surface in N=4 SU(N) SYM theory using localization on the four-dimensional ellipsoid. We extract the l…