Learning Transferable Visual Models From Natural Language Supervision
arXiv:2103.00020
Abstract
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept. Learning directly from raw text about images is a promising alternative which leverages a much broader source of supervision. We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. After pre-training, natural language is used to reference learned visual concepts (or describe new ones) enabling zero-shot transfer of the model to downstream tasks. We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, action recognition in videos, geo-localization, and many types of fine-grained object classification. The model transfers non-trivially to most tasks and is often competitive with a fully supervised baseline without the need for any dataset specific training. For instance, we match the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on. We release our code and pre-trained model weights at https://github.com/OpenAI/CLIP.
References in corpus (23)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Distributed Representations of Sentences and Documents
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Deep Learning Scaling is Predictable, Empirically
- Do ImageNet Classifiers Generalize to ImageNet?
- Measuring Robustness to Natural Distribution Shifts in Image Classification
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Learning and Evaluating General Linguistic Intelligence
- ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
- Jukebox: A Generative Model for Music
- Deep Structured Output Learning for Unconstrained Text Recognition
- Generating Wikipedia by Summarizing Long Sequences
- The Effect of Natural Distribution Shift on Question Answering Models
- TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
- RareAct: A video dataset of unusual interactions
- Exposing and Correcting the Gender Bias in Image Captioning Datasets and Models
Cited by in corpus (105)
- Diffusion Models Beat GANs on Image Synthesis
- Evaluating Large Language Models Trained on Code
- Zero-Shot Text-to-Image Generation
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- True Few-Shot Learning with Language Models
- ActionCLIP: A New Paradigm for Video Action Recognition
- StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
- Exploring the Limits of Out-of-Distribution Detection
- CLIPort: What and Where Pathways for Robotic Manipulation
- PanGu-: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
- CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders
- Multimodal Few-Shot Learning with Frozen Language Models
- A Simple Fix to Mahalanobis Distance for Improving Near-OOD Detection
- Partial success in closing the gap between human and machine vision
- Multiscale Vision Transformers
- MERLOT: Multimodal Neural Script Knowledge Models
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- What Makes Multi-modal Learning Better than Single (Provably)
- Container: Context Aggregation Network
- VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
- Unifying Multimodal Transformer for Bi-directional Image and Text Generation
- Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications
- Exploring the Limits of Large Scale Pre-training
- VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
- Neural Symbolic Regression that Scales
- Scaling Laws for Transfer
- LanguageRefer: Spatial-Language Model for 3D Visual Grounding
- Towards Open-World Text-Guided Face Image Generation and Manipulation
- The Evolution of Out-of-Distribution Robustness Throughout Fine-Tuning
- ClipMatrix: Text-controlled Creation of 3D Textured Meshes
- Language Grounding with 3D Objects
- Multi-Label Image Classification with Contrastive Learning
- MixNorm: Test-Time Adaptation Through Online Normalization Estimation
- Object-aware Contrastive Learning for Debiased Scene Representation
- Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
- MURAL: Multimodal, Multitask Retrieval Across Languages
- Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation
- Differentiable Physics: A Position Piece
- Temperature as Uncertainty in Contrastive Learning
- FairyTailor: A Multimodal Generative Framework for Storytelling
- Simpler, Faster, Stronger: Breaking The log-K Curse On Contrastive Learners With FlatNCE
- Large-Scale Zero-Shot Image Classification from Rich and Diverse Textual Descriptions
- Active Divergence with Generative Deep Learning -- A Survey and Taxonomy
- EfficientCLIP: Efficient Cross-Modal Pre-training by Ensemble Confident Learning and Language Modeling
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- A CLIP-Enhanced Method for Video-Language Understanding
- LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval
- Positive Sample Propagation along the Audio-Visual Event Line
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- An Image is Worth More Than a Thousand Words: Towards Disentanglement in the Wild
- On the Expressive Power of Self-Attention Matrices
- Paint4Poem: A Dataset for Artistic Visualization of Classical Chinese Poems
- SoK: How Robust is Image Classification Deep Neural Network Watermarking? (Extended Version)
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- CLIP4Caption ++: Multi-CLIP for Video Caption
- Automating Generative Deep Learning for Artistic Purposes: Challenges and Opportunities
- Winning the ICCV'2021 VALUE Challenge: Task-aware Ensemble and Transfer Learning with Visual Concepts
- Is Object Detection Necessary for Human-Object Interaction Recognition?
- Visual Conceptual Blending with Large-scale Language and Vision Models
- BEAMetrics: A Benchmark for Language Generation Evaluation Evaluation
- Connecting Language and Vision for Natural Language-Based Vehicle Retrieval
- Auditing AI models for Verified Deployment under Semantic Specifications
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer
- EncoderMI: Membership Inference against Pre-trained Encoders in Contrastive Learning
- Evolving Evocative 2D Views of Generated 3D Objects
- Telling Creative Stories Using Generative Visual Aids
- Less is More: Sparse Sampling for Dense Reaction Predictions
- ViSeRet: A simple yet effective approach to moment retrieval via fine-grained video segmentation
- Pretrained Encoders are All You Need
- Integrating Auxiliary Information in Self-supervised Learning
- Data-Efficient Language-Supervised Zero-Shot Learning with Self-Distillation
- Zero-Shot Information Extraction as a Unified Text-to-Triple Translation
- Does Vision-and-Language Pretraining Improve Lexical Grounding?
- EVOQUER: Enhancing Temporal Grounding with Video-Pivoted BackQuery Generation
- What Matters for Ad-hoc Video Search? A Large-scale Evaluation on TRECVID
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-modal Pretraining
- Will Multi-modal Data Improves Few-shot Learning?
- eProduct: A Million-Scale Visual Search Benchmark to Address Product Recognition Challenges
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Planning Multimodal Exploratory Actions for Online Robot Attribute Learning
- Personalizing Pre-trained Models
- Reading Isn't Believing: Adversarial Attacks On Multi-Modal Neurons
- Contrastive Fine-tuning Improves Robustness for Neural Rankers
- Evo* 2020 -- Late-Breaking Abstracts Volume
- Artistic Autonomy in AI Art
- Scaling Laws for the Few-Shot Adaptation of Pre-trained Image Classifiers
- Embed Everything: A Method for Efficiently Co-Embedding Multi-Modal Spaces
- Language Models as Zero-shot Visual Semantic Learners
- CLIP4Caption: CLIP for Video Caption
- Fine-Grained Chemical Entity Typing with Multimodal Knowledge Representation
- An animated picture says at least a thousand words: Selecting Gif-based Replies in Multimodal Dialog
- Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast
- Data Efficient Masked Language Modeling for Vision and Language
- Few-Shot Intent Detection via Contrastive Pre-Training and Fine-Tuning
- Inverse Problems Leveraging Pre-trained Contrastive Representations
- Cross-Modal Retrieval Augmentation for Multi-Modal Classification
- Cut the CARP: Fishing for zero-shot story evaluation
- Unsupervised Source Separation By Steering Pretrained Music Models
- MOMENTA: A Multimodal Framework for Detecting Harmful Memes and Their Targets