Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
arXiv:2102.05918
Abstract
Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language representations still rely heavily on curated training datasets that are expensive or require expert knowledge. For vision applications, representations are mostly learned using datasets with explicit class labels such as ImageNet or OpenImages. For vision-language, popular datasets like Conceptual Captions, MSCOCO, or CLIP all involve a non-trivial data collection (and cleaning) process. This costly curation process limits the size of datasets and hence hinders the scaling of trained models. In this paper, we leverage a noisy dataset of over one billion image alt-text pairs, obtained without expensive filtering or post-processing steps in the Conceptual Captions dataset. A simple dual-encoder architecture learns to align visual and language representations of the image and text pairs using a contrastive loss. We show that the scale of our corpus can make up for its noise and leads to state-of-the-art representations even with such a simple learning scheme. Our visual representation achieves strong performance when transferred to classification tasks such as ImageNet and VTAB. The aligned visual and language representations enables zero-shot image classification and also set new state-of-the-art results on Flickr30K and MSCOCO image-text retrieval benchmarks, even when compared with more sophisticated cross-attention models. The representations also enable cross-modality search with complex text and text + image queries.
ICML 2021
References in corpus (21)
- Efficient Estimation of Word Representations in Vector Space
- Distributed Representations of Words and Phrases and their Compositionality
- Learning Transferable Visual Models From Natural Language Supervision
- Bootstrap your own latent: A new approach to self-supervised Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Do ImageNet Classifiers Generalize to ImageNet?
- Billion-scale semi-supervised learning for image classification
- Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
- M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training
- Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations
- Graph-RISE: Graph-Regularized Image Semantic Embedding
- Learning the Best Pooling Strategy for Visual Semantic Embedding
- UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training
Cited by in corpus (114)
- Learning to Prompt for Vision-Language Models
- MLP-Mixer: An all-MLP Architecture for Vision
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Reproducible scaling laws for contrastive language-image learning
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Florence: A New Foundation Model for Computer Vision
- Compute Trends Across Three Eras of Machine Learning
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
- Towards artificial general intelligence via a multimodal foundation model
- Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
- VLP: A Survey on Vision-Language Pre-training
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
- A Foundation Language-Image Model of the Retina (FLAIR): Encoding Expert Knowledge in Text Supervision
- Multimodal Few-Shot Learning with Frozen Language Models
- CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training
- Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
- Efficient Token-Guided Image-Text Retrieval with Consistent Multimodal Contrastive Training
- Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
- LAFITE: Towards Language-Free Training for Text-to-Image Generation
- Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications
- Several categories of Large Language Models (LLMs): A Short Survey
- How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose
- UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
- ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust Learning
- CgT-GAN: CLIP-guided Text GAN for Image Captioning
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages
- An Empirical Study of Training End-to-End Vision-and-Language Transformers
- Towards Open-World Text-Guided Face Image Generation and Manipulation
- Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
- Poisoning and Backdooring Contrastive Learning
- Breaking with Fixed Set Pathology Recognition through Report-Guided Contrastive Training
- Contrastive Attention for Automatic Chest X-ray Report Generation
- Zero-Shot Text-Guided Object Generation with Dream Fields
- Scaling Up Vision-Language Pre-training for Image Captioning
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- Hierarchical Matching and Reasoning for Multi-Query Image Retrieval
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Joint Learning of Localized Representations from Medical Images and Reports
- Homogeneous vector bundles and -equivariant convolutional neural networks
- MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning
- MURAL: Multimodal, Multitask Retrieval Across Languages
- Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning
- Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation
- A Simple Long-Tailed Recognition Baseline via Vision-Language Model
- Fusing Domain-Specific Content from Large Language Models into Knowledge Graphs for Enhanced Zero Shot Object State Classification
- INTERN: A New Learning Paradigm Towards General Vision
- Simpler, Faster, Stronger: Breaking The log-K Curse On Contrastive Learners With FlatNCE
- Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
- Generative Art Using Neural Visual Grammars and Dual Encoders
- Synthetic Boost: Leveraging Synthetic Data for Enhanced Vision-Language Segmentation in Echocardiography
- MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training
- EfficientCLIP: Efficient Cross-Modal Pre-training by Ensemble Confident Learning and Language Modeling
- Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data
- Semantically Guided Representation Learning For Action Anticipation
- Positive Sample Propagation along the Audio-Visual Event Line
- Image-Text Pre-Training for Logo Recognition
- Multi-Task Self-Training for Learning General Representations
- Don't waste SAM
- MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
- ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
- Multi-label Cluster Discrimination for Visual Representation Learning
- AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation
- MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
- On the Expressive Power of Self-Attention Matrices
- No One Representation to Rule Them All: Overlapping Features of Training Methods
- Robustness Tokens: Towards Adversarial Robustness of Transformers
- Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
- Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation
- 3D Object Detection and High-Resolution Traffic Parameters Extraction Using Low-Resolution LiDAR Data
- Exploring the Transferability of a Foundation Model for Fundus Images: Application to Hypertensive Retinopathy
- Centered Masking for Language-Image Pre-Training
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- Data-Efficient Language-Supervised Zero-Shot Learning with Self-Distillation
- Dual-Modality Representation Learning for Molecular Property Prediction
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- Achieving Human Parity on Visual Question Answering
- VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition
- You Never Cluster Alone
- Goal-driven text descriptions for images
- PØDA: Prompt-driven Zero-shot Domain Adaptation
- Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment
- TPA3D: Triplane Attention for Fast Text-to-3D Generation
- CARLS: Cross-platform Asynchronous Representation Learning System
- Collaborating Foundation Models for Domain Generalized Semantic Segmentation
- TransAug: Translate as Augmentation for Sentence Embeddings
- GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
- Continually Learn to Map Visual Concepts to Large Language Models in Resource-constrained Environments
- Temporal Image Caption Retrieval Competition -- Description and Results
- EEG-Driven Image Reconstruction with Saliency-Guided Diffusion Models
- Visual Objectification in Films: Towards a New AI Task for Video Interpretation
- Cross-Modal Retrieval Augmentation for Multi-Modal Classification
- Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval