108 citations · 272 across the 7 of their papers we have counts for
6 papers · 1 filter
When Do Curricula Work?
Xiaoxia Wu, Ethan Dyer, Behnam Neyshabur
Inspired by human learning, researchers have proposed ordering examples during training based on their difficulty. Both curriculum learning, exposing a network to easier examples e…
Asymptotics of Wide Convolutional Neural Networks
Anders Andreassen, Ethan Dyer
Wide neural networks have proven to be a rich class of architectures for both theory and practice. Motivated by the observation that finite width convolutional networks appear to o…
Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics
Vinay V. Ramasesh, Ethan Dyer, Maithra Raghu
A central challenge in developing versatile machine learning systems is catastrophic forgetting: a model trained on tasks in sequence will suffer significant performance drops on e…
Affinity and Diversity: Quantifying Mechanisms of Data Augmentation
Raphael Gontijo-Lopes, Sylvia J. Smullin, Ekin D. Cubuk +1
Though data augmentation has become a standard component of deep neural network training, the underlying mechanism behind the effectiveness of these techniques remains poorly under…
Asymptotics of Wide Networks from Feynman Diagrams
Ethan Dyer, Guy Gur-Ari
Understanding the asymptotic behavior of wide networks is of considerable interest. In this work, we present a general method for analyzing this large width behavior. The method is…
Gradient Descent Happens in a Tiny Subspace
Guy Gur-Ari, Daniel A. Roberts, Ethan Dyer
We show that in a variety of large-scale deep learning scenarios the gradient dynamically converges to a very small subspace after a short period of training. The subspace is spann…