SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
arXiv:2108.10904
Abstract
With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations including clean image captions and regional labels limits the scalability of existing approaches, and complicates the pretraining procedure with the introduction of multiple dataset-specific objectives. In this work, we relax these constraints and present a minimalist pretraining framework, named Simple Visual Language Model (SimVLM). Unlike prior work, SimVLM reduces the training complexity by exploiting large-scale weak supervision, and is trained end-to-end with a single prefix language modeling objective. Without utilizing extra data or task-specific customization, the resulting model significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA (+3.74% vqa-score), NLVR2 (+1.17% accuracy), SNLI-VE (+1.37% accuracy) and image captioning tasks (+10.1% average CIDEr score). Furthermore, we demonstrate that SimVLM acquires strong generalization and transfer ability, enabling zero-shot behavior including open-ended visual question answering and cross-modality transfer.
Published at ICLR 2022
References in corpus (8)
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Cross-lingual Language Model Pretraining
- Zero-Shot Text-to-Image Generation
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Cited by in corpus (31)
- Florence: A New Foundation Model for Computer Vision
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
- VLP: A Survey on Vision-Language Pre-training
- BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Exploring scalable medical image encoders beyond text supervision
- OpenViVQA: Task, Dataset, and Multimodal Fusion Models for Visual Question Answering in Vietnamese
- CgT-GAN: CLIP-guided Text GAN for Image Captioning
- Learning to Evaluate Performance of Multi-modal Semantic Localization
- An Empirical Study of Training End-to-End Vision-and-Language Transformers
- Scaling Up Vision-Language Pre-training for Image Captioning
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- MMSR: Symbolic Regression is a Multi-Modal Information Fusion Task
- Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
- Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations
- ClipCap: CLIP Prefix for Image Captioning
- MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation
- Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data
- STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-training
- An Empirical Study on the Language Modal in Visual Question Answering
- AI as a Tool for Fair Journalism: Case Studies from Malta
- Human Inspired Progressive Alignment and Comparative Learning for Grounded Word Acquisition
- Top1 Solution of QQ Browser 2021 Ai Algorithm Competition Track 1 : Multimodal Video Similarity
- Convolutional Gated MLP: Combining Convolutions & gMLP
- Achieving Human Parity on Visual Question Answering
- Less is More: Generating Grounded Navigation Instructions from Landmarks
- Partially-Supervised Novel Object Captioning Leveraging Context from Paired Data