7 citations · 7 across the 7 of their papers we have counts for
7 papers · 1 filter
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
Mohammadreza Salehi, Shashanka Venkataramanan, Ioana Simion +3
Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of…
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi +5
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance…
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
Florent Bartoccioni, Elias Ramzi, Victor Besnier +14
We explore the potential of large-scale generative video models for autonomous driving, introducing an open-source auto-regressive video model (VaViM) and its companion video-actio…
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
Shashanka Venkataramanan, Mamshad Nayeem Rizve, João Carreira +2
Self-supervised learning has unlocked the potential of scaling up pretraining to billions of images, since annotation is unnecessary. But are we making the best use of data? How mo…
Skip-Attention: Improving Vision Transformers by Paying Less Attention
Shashanka Venkataramanan, Amir Ghodrati, Yuki M. Asano +2
This work aims to improve the efficiency of vision transformers (ViT). While ViTs use computationally expensive self-attention operations in every layer, we identify that these ope…
AlignMixup: Improving Representations By Interpolating Aligned Features
Shashanka Venkataramanan, Ewa Kijak, Laurent Amsaleg +1
Mixup is a powerful data augmentation method that interpolates between two or more examples in the input or feature space and between the corresponding target labels. Many recent m…