26 citations · 45 across the 4 of their papers we have counts for
4 papers · 1 filter
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
Revant Teotia, Adrien Bardes, Michael Rabbat +3
Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in vid…
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
Quentin Garrido, Nicolas Ballas, Mahmoud Assran +5
We investigate the emergence of intuitive physics understanding in general-purpose deep neural network models trained to predict masked regions in natural videos. Leveraging the vi…
Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks
Micah Goldblum, Hossein Souri, Renkun Ni +10
Neural network based computer vision systems are typically built on a backbone, a pretrained or randomly initialized feature extractor. Several years ago, the default option was an…
VICRegL: Self-Supervised Learning of Local Visual Features
Adrien Bardes, Jean Ponce, Yann LeCun
Most recent self-supervised methods for learning image representations focus on either producing a global feature with invariance properties, or producing a set of local features.…