VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
arXiv:2111.02358
Abstract
We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Mixture-of-Modality-Experts (MoME) Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of MoME, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. The code and pretrained models are available at https://aka.ms/vlmo.
Work in progress
References in corpus (9)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- BEiT: BERT Pre-Training of Image Transformers
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- DeltaLM: Encoder-Decoder Pre-training for Language Generation and Translation by Augmenting Pretrained Multilingual Encoders
Cited by in corpus (14)
- VLP: A Survey on Vision-Language Pre-training
- RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
- Cross-modal Contrastive Learning for Multimodal Fake News Detection
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
- GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
- Component Segmentation of Engineering Drawings Using Graph Convolutional Networks
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
- Multimodal Neural Databases
- Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
- MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
- Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
- sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging