Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
arXiv:2110.05208
Abstract
Recently, large-scale Contrastive Language-Image Pre-training (CLIP) has attracted unprecedented attention for its impressive zero-shot recognition ability and excellent transferability to downstream tasks. However, CLIP is quite data-hungry and requires 400M image-text pairs for pre-training, thereby restricting its adoption. This work proposes a novel training paradigm, Data efficient CLIP (DeCLIP), to alleviate this limitation. We demonstrate that by carefully utilizing the widespread supervision among the image-text pairs, our De-CLIP can learn generic visual features more efficiently. Instead of using the single image-text contrastive supervision, we fully exploit data potential through the use of (1) self-supervision within each modality; (2) multi-view supervision across modalities; (3) nearest-neighbor supervision from other similar pairs. Benefiting from intrinsic supervision, our DeCLIP-ResNet50 can achieve 60.4% zero-shot top1 accuracy on ImageNet, which is 0.8% above the CLIP-ResNet50 while using 7.1 x fewer data. Our DeCLIP-ResNet50 outperforms its counterpart in 8 out of 11 visual datasets when transferred to downstream tasks. Moreover, Scaling up the model and computing also works well in our framework.Our code, dataset and models are released at: https://github.com/Sense-GVT/DeCLIP
17 pages, 10 figures
References in corpus (18)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Simple Framework for Contrastive Learning of Visual Representations
- Learning Transferable Visual Models From Natural Language Supervision
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- YFCC100M: The New Data in Multimedia Research
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- LXMERT: Learning Cross-Modality Encoder Representations from Transformers
- EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- Self-supervised Pretraining of Visual Features in the Wild
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- Revisiting Contrastive Methods for Unsupervised Learning of Visual Representations
Cited by in corpus (8)
- VLP: A Survey on Vision-Language Pre-training
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- Exploring scalable medical image encoders beyond text supervision
- CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
- How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose
- Synthetic Boost: Leveraging Synthetic Data for Enhanced Vision-Language Segmentation in Echocardiography
- Gen-AI for User Safety: A Survey