50 citations · 208 across the 12 of their papers we have counts for
31 papers
lo-fi: distributed fine-tuning without communication
Mitchell Wortsman, Suchin Gururangan, Shen Li +4
When fine-tuning large neural networks, it is common to use multiple nodes and to communicate gradients at each optimization step. By contrast, we investigate completely local fine…
The Robustness Limits of SoTA Vision Models to Natural Variation
Mark Ibrahim, Quentin Garrido, Ari Morcos +1
Recent state-of-the-art vision models introduced new architectures, learning paradigms, and larger pretraining data, leading to impressive performance on tasks such as classificati…
Robust Self-Supervised Learning with Lie Groups
Mark Ibrahim, Diane Bouchacourt, Ari Morcos
Deep learning has led to remarkable advances in computer vision. Even so, today's best models are brittle when presented with variations that differ even slightly from those seen d…
Trade-offs of Local SGD at Scale: An Empirical Study
Jose Javier Gonzalez Ortiz, Jonathan Frankle, Mike Rabbat +2
As datasets and models become increasingly large, distributed training has become a necessary component to allow deep neural networks to train in reasonable amounts of time. Howeve…
Transformed CNNs: recasting pre-trained convolutional layers with self-attention
Stéphane d'Ascoli, Levent Sagun, Giulio Biroli +1
Vision Transformers (ViT) have recently emerged as a powerful alternative to convolutional networks (CNNs). Although hybrid models attempt to bridge the gap between these two archi…
Width Transfer: On the (In)variance of Width Optimization
Ting-Wu Chin, Diana Marculescu, Ari S. Morcos
Optimizing the channel counts for different layers of a CNN has shown great promise in improving the efficiency of CNNs at test-time. However, these methods often introduce large c…