iBOT: Image BERT Pre-Training with Online Tokenizer
arXiv:2111.07832
Abstract
The success of language Transformers is primarily attributed to the pretext task of masked language modeling (MLM), where texts are first tokenized into semantically meaningful pieces. In this work, we study masked image modeling (MIM) and indicate the advantages and challenges of using a semantically meaningful visual tokenizer. We present a self-supervised framework iBOT that can perform masked prediction with an online tokenizer. Specifically, we perform self-distillation on masked patch tokens and take the teacher network as the online tokenizer, along with self-distillation on the class token to acquire visual semantics. The online tokenizer is jointly learnable with the MIM objective and dispenses with a multi-stage training pipeline where the tokenizer needs to be pre-trained beforehand. We show the prominence of iBOT by achieving an 82.3% linear probing accuracy and an 87.8% fine-tuning accuracy evaluated on ImageNet-1K. Beyond the state-of-the-art image classification results, we underline emerging local semantic patterns, which helps the models to obtain strong robustness against common corruptions and achieve leading results on dense downstream tasks, eg., object detection, instance segmentation, and semantic segmentation.
Cited by in corpus (20)
- DINOv2: Learning Robust Visual Features without Supervision
- CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding
- Vision Transformers, a new approach for high-resolution and large-scale mapping of canopy heights
- AstroCLIP: A Cross-Modal Foundation Model for Galaxies
- Vision Transformers Need Registers
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- Enhancing Network Initialization for Medical AI Models Using Large-Scale, Unlabeled Natural Images
- Evaluating Pre-trained Convolutional Neural Networks and Foundation Models as Feature Extractors for Content-based Medical Image Retrieval
- Self-supervised visual learning in the low-data regime: a comparative evaluation
- MambaMIM: Pre-training Mamba with State Space Token Interpolation and its Application to Medical Image Segmentation
- Exploring Advances in Transformers and CNN for Skin Lesion Diagnosis on Small Datasets
- In-Domain Self-Supervised Learning Improves Remote Sensing Image Scene Classification
- Boosting multi-demographic federated learning for chest radiograph analysis using general-purpose self-supervised representations
- DatUS^2: Data-driven Unsupervised Semantic Segmentation with Pre-trained Self-supervised Vision Transformer
- Self-supervised Model Based on Masked Autoencoders Advance CT Scans Classification
- Astronomical image time series classification using CONVolutional attENTION (ConvEntion)
- Leveraging Self-Supervised Vision Transformers for Segmentation-based Transfer Function Design
- SERE: Exploring Feature Self-relation for Self-supervised Transformer
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- jBOT: Semantic Jet Representation Clustering Emerges from Self-Distillation