BEiT: BERT Pre-Training of Image Transformers
arXiv:2106.08254
Abstract
We introduce a self-supervised vision representation model BEiT, which stands for Bidirectional Encoder representation from Image Transformers. Following BERT developed in the natural language processing area, we propose a masked image modeling task to pretrain vision Transformers. Specifically, each image has two views in our pre-training, i.e, image patches (such as 16x16 pixels), and visual tokens (i.e., discrete tokens). We first "tokenize" the original image into visual tokens. Then we randomly mask some image patches and fed them into the backbone Transformer. The pre-training objective is to recover the original visual tokens based on the corrupted image patches. After pre-training BEiT, we directly fine-tune the model parameters on downstream tasks by appending task layers upon the pretrained encoder. Experimental results on image classification and semantic segmentation show that our model achieves competitive results with previous pre-training methods. For example, base-size BEiT achieves 83.2% top-1 accuracy on ImageNet-1K, significantly outperforming from-scratch DeiT training (81.8%) with the same setup. Moreover, large-size BEiT obtains 86.3% only using ImageNet-1K, even outperforming ViT-L with supervised pre-training on ImageNet-22K (85.2%). The code and pretrained models are available at https://aka.ms/beit.
A Path to the BERT Moment of CV
References in corpus (5)
Cited by in corpus (77)
- Attention Mechanisms in Computer Vision: A Survey
- Evaluating Large Language Models Trained on Code
- DINOv2: Learning Robust Visual Features without Supervision
- Reproducible scaling laws for contrastive language-image learning
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
- TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting
- Efficient Multimodal Transformer with Dual-Level Feature Restoration for Robust Multimodal Sentiment Analysis
- UNETR: Transformers for 3D Medical Image Segmentation
- Masked Autoencoders Are Scalable Vision Learners
- Label-Efficient Self-Supervised Federated Learning for Tackling Data Heterogeneity in Medical Imaging
- Contrastive Masked Autoencoders are Stronger Vision Learners
- Improving Chest X-Ray Report Generation by Leveraging Warm Starting
- SignBERT+: Hand-model-aware Self-supervised Pre-training for Sign Language Understanding
- Swin Transformer V2: Scaling Up Capacity and Resolution
- Masked-attention Mask Transformer for Universal Image Segmentation
- Advances of Machine Learning in Materials Science: Ideas and Techniques
- CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding
- Vector-quantized Image Modeling with Improved VQGAN
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- HTR-VT: Handwritten Text Recognition with Vision Transformer
- Benchmarking Detection Transfer Learning with Vision Transformers
- TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models
- Towards Label-efficient Automatic Diagnosis and Analysis: A Comprehensive Survey of Advanced Deep Learning-based Weakly-supervised, Semi-supervised and Self-supervised Techniques in Histopathological Image Analysis
- Multi-scale Transformer Network with Edge-aware Pre-training for Cross-Modality MR Image Synthesis
- Vision Transformers Need Registers
- Do We Really Need Dice? The Hidden Region-Size Biases of Segmentation Losses
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training
- A Survey of Visual Transformers
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- SVFAP: Self-supervised Video Facial Affect Perceiver
- SelfFed: Self-Supervised Federated Learning for Data Heterogeneity and Label Scarcity in Medical Images
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Deep Active Learning for Computer Vision: Past and Future
- Re-Scoring Using Image-Language Similarity for Few-Shot Object Detection
- MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping
- Self-supervised visual learning in the low-data regime: a comparative evaluation
- Less is More: Consistent Video Depth Estimation with Masked Frames Modeling
- An Empirical Study of Training End-to-End Vision-and-Language Transformers
- MambaMIM: Pre-training Mamba with State Space Token Interpolation and its Application to Medical Image Segmentation
- EDMAE: An Efficient Decoupled Masked Autoencoder for Standard View Identification in Pediatric Echocardiography
- Exploring Advances in Transformers and CNN for Skin Lesion Diagnosis on Small Datasets
- EVJVQA Challenge: Multilingual Visual Question Answering
- Dilated convolution with learnable spacings
- Language Modelling with Pixels
- SimMIM: A Simple Framework for Masked Image Modeling
- Robust Lane Detection through Self Pre-training with Masked Sequential Autoencoders and Fine-tuning with Customized PolyLoss
- Low-resource finetuning of foundation models beats state-of-the-art in histopathology
- Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning
- Multi-modal Medical Image Fusion For Non-Small Cell Lung Cancer Classification
- Generating 3D Bio-Printable Patches Using Wound Segmentation and Reconstruction to Treat Diabetic Foot Ulcers
- INTERN: A New Learning Paradigm Towards General Vision
- Discrete Representations Strengthen Vision Transformer Robustness
- Rethinking Supervised Pre-training for Better Downstream Transferring
- Modularity in Deep Learning: A Survey
- On The State of Data In Computer Vision: Human Annotations Remain Indispensable for Developing Deep Learning Models
- Semantic-Aware Generation for Self-Supervised Visual Representation Learning
- Self-supervised Semi-supervised Learning for Data Labeling and Quality Evaluation
- Co-Salient Object Detection with Semantic-Level Consensus Extraction and Dispersion
- Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
- Pelta: Shielding Transformers to Mitigate Evasion Attacks in Federated Learning
- Pyramid Adversarial Training Improves ViT Performance
- Toward High Quality Facial Representation Learning
- Impact of Ground Truth Quality on Handwriting Recognition
- Multi-modal Document Presentation Attack Detection With Forensics Trace Disentanglement
- Downstream-Pretext Domain Knowledge Traceback for Active Learning
- 3rd Place Solution for VisDA 2021 Challenge -- Universally Domain Adaptive Image Recognition
- Leveraging Self-Supervised Vision Transformers for Segmentation-based Transfer Function Design
- On Masked Pre-training and the Marginal Likelihood
- TransMix: Attend to Mix for Vision Transformers
- One to Transfer All: A Universal Transfer Framework for Vision Foundation Model with Few Data
- Memory Based Video Scene Parsing
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- SERE: Exploring Feature Self-relation for Self-supervised Transformer
- EdiBERT, a generative model for image editing
- Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners
- MC-SSL0.0: Towards Multi-Concept Self-Supervised Learning