1.4k citations · 2.6k across the 12 of their papers we have counts for
7 papers · 1 filter
MLP-Mixer: An all-MLP Architecture for Vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov +9
Convolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this…
SI-Score: An image dataset for fine-grained analysis of robustness to object location, rotation and size
Jessica Yung, Rob Romijnders, Alexander Kolesnikov +6
Before deploying machine learning models it is critical to assess their robustness. In the context of deep neural networks for image understanding, changing the object location, ro…
ViViT: A Video Vision Transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold +3
We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatio-temporal tokens…
Representation learning from videos in-the-wild: An object-centric approach
Rob Romijnders, Aravindh Mahendran, Michael Tschannen +4
We propose a method to learn image representations from uncurated videos. We combine a supervised loss from off-the-shelf object detectors and self-supervised losses which naturall…
On Robustness and Transferability of Convolutional Neural Networks
Josip Djolonga, Jessica Yung, Michael Tschannen +11
Modern deep convolutional networks (CNNs) are often criticized for not generalizing under distributional shifts. However, several recent breakthroughs in transfer learning suggest…
Self-Supervised Learning of Video-Induced Visual Invariances
Michael Tschannen, Josip Djolonga, Marvin Ritter +5
We propose a general framework for self-supervised learning of transferable visual representations based on Video-Induced Visual Invariances (VIVI). We consider the implicit hierar…