SERE: Exploring Feature Self-relation for Self-supervised Transformer
arXiv:2206.05184
Abstract
Learning representations with self-supervision for convolutional networks (CNN) has been validated to be effective for vision tasks. As an alternative to CNN, vision transformers (ViT) have strong representation ability with spatial self-attention and channel-level feedforward networks. Recent works reveal that self-supervised learning helps unleash the great potential of ViT. Still, most works follow self-supervised strategies designed for CNN, e.g., instance-level discrimination of samples, but they ignore the properties of ViT. We observe that relational modeling on spatial and channel dimensions distinguishes ViT from other networks. To enforce this property, we explore the feature SElf-RElation (SERE) for training self-supervised ViT. Specifically, instead of conducting self-supervised learning solely on feature embeddings from multiple views, we utilize the feature self-relations, i.e., spatial/channel self-relations, for self-supervised learning. Self-relation based learning further enhances the relation modeling ability of ViT, resulting in stronger representations that stably improve performance on multiple downstream tasks. Our source code is publicly available at: https://github.com/MCG-NKU/SERE.
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
References in corpus (23)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Simple Framework for Contrastive Learning of Visual Representations
- Bootstrap your own latent: A new approach to self-supervised Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- BEiT: BERT Pre-Training of Image Transformers
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- P2T: Pyramid Pooling Transformer for Scene Understanding
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
- iBOT: Image BERT Pre-Training with Online Tokenizer
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset Biases
- Unsupervised Semantic Segmentation by Distilling Feature Correspondences
- Self-Supervised Learning with Swin Transformers
- Localizing Objects with Self-Supervised Transformers and no Labels
- What to Hide from Your Students: Attention-Guided Masked Image Modeling
- Large-scale Unsupervised Semantic Segmentation
- Efficient Self-supervised Vision Transformers for Representation Learning
- SiT: Self-supervised vIsion Transformer
- A Review of Predictive and Contrastive Self-supervised Learning for Medical Images
- Mugs: A Multi-Granular Self-Supervised Learning Framework
- Revitalizing CNN Attentions via Transformers in Self-Supervised Visual Representation Learning
- Towards Sustainable Self-supervised Learning