activity
20182022
most citedThe Early Phase of Neural Network Training

50 citations · 208 across the 12 of their papers we have counts for

collaborators

31 papers

cs.LG20227 cited

lo-fi: distributed fine-tuning without communication

Mitchell Wortsman, Suchin Gururangan, Shen Li +4

When fine-tuning large neural networks, it is common to use multiple nodes and to communicate gradients at each optimization step. By contrast, we investigate completely local fine…

cs.CV20229 cited

The Robustness Limits of SoTA Vision Models to Natural Variation

Mark Ibrahim, Quentin Garrido, Ari Morcos +1

Recent state-of-the-art vision models introduced new architectures, learning paradigms, and larger pretraining data, leading to impressive performance on tasks such as classificati…

cs.CV20222 cited

Robust Self-Supervised Learning with Lie Groups

Mark Ibrahim, Diane Bouchacourt, Ari Morcos

Deep learning has led to remarkable advances in computer vision. Even so, today's best models are brittle when presented with variations that differ even slightly from those seen d…

cs.LG2021

Trade-offs of Local SGD at Scale: An Empirical Study

Jose Javier Gonzalez Ortiz, Jonathan Frankle, Mike Rabbat +2

As datasets and models become increasingly large, distributed training has become a necessary component to allow deep neural networks to train in reasonable amounts of time. Howeve…

cs.LG20213 cited

Transformed CNNs: recasting pre-trained convolutional layers with self-attention

Stéphane d'Ascoli, Levent Sagun, Giulio Biroli +1

Vision Transformers (ViT) have recently emerged as a powerful alternative to convolutional networks (CNNs). Although hybrid models attempt to bridge the gap between these two archi…

cs.CV2021

Width Transfer: On the (In)variance of Width Optimization

Ting-Wu Chin, Diana Marculescu, Ari S. Morcos

Optimizing the channel counts for different layers of a CNN has shown great promise in improving the efficiency of CNNs at test-time. However, these methods often introduce large c…