1.4k citations · 1.8k across the 11 of their papers we have counts for
16 papers
Better plain ViT baselines for ImageNet-1k
Lucas Beyer, Xiaohua Zhai, Alexander Kolesnikov
It is commonly accepted that the Vision Transformer model requires sophisticated regularization techniques to excel at ImageNet-1k scale data. Surprisingly, we find this is not the…
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer +32
Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and trainin…
MLP-Mixer: An all-MLP Architecture for Vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov +9
Convolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this…
SI-Score: An image dataset for fine-grained analysis of robustness to object location, rotation and size
Jessica Yung, Rob Romijnders, Alexander Kolesnikov +6
Before deploying machine learning models it is critical to assess their robustness. In the context of deep neural networks for image understanding, changing the object location, ro…
On Robustness and Transferability of Convolutional Neural Networks
Josip Djolonga, Jessica Yung, Michael Tschannen +11
Modern deep convolutional networks (CNNs) are often criticized for not generalizing under distributional shifts. However, several recent breakthroughs in transfer learning suggest…
Are we done with ImageNet?
Lucas Beyer, Olivier J. Hénaff, Alexander Kolesnikov +2
Yes, and no. We ask whether recent progress on the ImageNet classification benchmark continues to represent meaningful generalization, or whether the community has started to overf…