PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
arXiv:2111.12710
Abstract
This paper explores a better prediction target for BERT pre-training of vision transformers. We observe that current prediction targets disagree with human perception judgment.This contradiction motivates us to learn a perceptual prediction target. We argue that perceptually similar images should stay close to each other in the prediction target space. We surprisingly find one simple yet effective idea: enforcing perceptual similarity during the dVAE training. Moreover, we adopt a self-supervised transformer model for deep feature extraction and show that it works well for calculating perceptual similarity.We demonstrate that such learned visual tokens indeed exhibit better semantic meanings, and help pre-training achieve superior transfer performance in various downstream tasks. For example, we achieve Top-1 accuracy on ImageNet-1K with ViT-B backbone, outperforming the competitive method BEiT by under the same pre-training epochs. Our approach also gets significant improvement on object detection and segmentation on COCO and semantic segmentation on ADE20K. Equipped with a larger backbone ViT-H, we achieve the state-of-the-art ImageNet accuracy (\textbf{88.3\%}) among methods using only ImageNet-1K data.
To appear at AAAI 2023
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- Zero-Shot Text-to-Image Generation
- BEiT: BERT Pre-Training of Image Transformers
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Learning Representations by Maximizing Mutual Information Across Views
- Stand-Alone Self-Attention in Vision Models
- Masked Autoencoders Are Scalable Vision Learners
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
Cited by in corpus (3)
- Contrastive Masked Autoencoders are Stronger Vision Learners
- CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding
- Towards Label-efficient Automatic Diagnosis and Analysis: A Comprehensive Survey of Advanced Deep Learning-based Weakly-supervised, Semi-supervised and Self-supervised Techniques in Histopathological Image Analysis