11 citations · 18 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 11 cited
Revisiting Feature Prediction for Learning Visual Representations from Video
Adrien Bardes, Quentin Garrido, Jean Ponce +5
This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a f…
cs.CV2024★ 7 cited
Learning and Leveraging World Models in Visual Representation Learning
Quentin Garrido, Mahmoud Assran, Nicolas Ballas +3
Joint-Embedding Predictive Architecture (JEPA) has emerged as a promising self-supervised approach that learns by leveraging a world model. While previously limited to predicting m…