Efficient Self-supervised Vision Transformers for Representation Learning
arXiv:2106.09785
Abstract
This paper investigates two techniques for developing efficient self-supervised vision transformers (EsViT) for visual representation learning. First, we show through a comprehensive empirical study that multi-stage architectures with sparse self-attentions can significantly reduce modeling complexity but with a cost of losing the ability to capture fine-grained correspondences between image regions. Second, we propose a new pre-training task of region matching which allows the model to capture fine-grained region dependencies and as a result significantly improves the quality of the learned vision representations. Our results show that combining the two techniques, EsViT achieves 81.3% top-1 on the ImageNet linear probe evaluation, outperforming prior arts with around an order magnitude of higher throughput. When transferring to downstream linear classification tasks, EsViT outperforms its supervised counterpart on 17 out of 18 datasets. The code and models are publicly available: https://github.com/microsoft/esvit
ICLR 2022; Code: https://github.com/microsoft/esvit
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Learning Transferable Visual Models From Natural Language Supervision
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- MLP-Mixer: An all-MLP Architecture for Vision
- Variational Autoencoder for Deep Learning of Images, Labels and Captions
- WebVision Database: Visual Learning and Understanding from Web Data
- CvT: Introducing Convolutions to Vision Transformers
- Self-supervised Pretraining of Visual Features in the Wild
- Pre-Trained Image Processing Transformer
- End-to-End Object Detection with Adaptive Clustering Transformer
- Self-Supervised Learning with Swin Transformers
- End-to-End Video Instance Segmentation with Transformers
- Self-supervised Pre-training with Hard Examples Improves Visual Representations
Cited by in corpus (12)
- A Survey on Visual Transformer
- Transformers in Vision: A Survey
- DINOv2: Learning Robust Visual Features without Supervision
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- Self-Supervised and Invariant Representations for Wireless Localization
- DimCL: Dimensional Contrastive Learning For Improving Self-Supervised Learning
- Self supervised learning improves dMMR/MSI detection from histology slides across multiple cancers
- Semantic-Aware Generation for Self-Supervised Visual Representation Learning
- SSIN: Self-Supervised Learning for Rainfall Spatial Interpolation
- Active Learning at the ImageNet Scale
- Characterizing and Improving the Robustness of Self-Supervised Learning through Background Augmentations
- SERE: Exploring Feature Self-relation for Self-supervised Transformer