WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
arXiv:2103.01913 · doi:10.1145/3404835.3463257
Abstract
The milestone improvements brought about by deep representation learning and pre-training techniques have led to large performance gains across downstream NLP, IR and Vision tasks. Multimodal modeling techniques aim to leverage large high-quality visio-linguistic datasets for learning complementary information (across image and text modalities). In this paper, we introduce the Wikipedia-based Image Text (WIT) Dataset (https://github.com/google-research-datasets/wit) to better facilitate multimodal, multilingual learning. WIT is composed of a curated set of 37.6 million entity rich image-text examples with 11.5 million unique images across 108 Wikipedia languages. Its size enables WIT to be used as a pretraining dataset for multimodal models, as we show when applied to downstream tasks such as image-text retrieval. WIT has four main and unique advantages. First, WIT is the largest multimodal dataset by the number of image-text examples by 3x (at the time of writing). Second, WIT is massively multilingual (first of its kind) with coverage over 100+ languages (each of which has at least 12K examples) and provides cross-lingual texts for many images. Third, WIT represents a more diverse set of concepts and real world entities relative to what previous datasets cover. Lastly, WIT provides a very challenging real-world test set, as we empirically illustrate using an image-text retrieval task as an example.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Learning Transferable Visual Models From Natural Language Supervision
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
- Graph-RISE: Graph-Regularized Image Semantic Embedding
Cited by in corpus (9)
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- Contrastive Language-Image Pre-training for the Italian Language
- Scaling Up Vision-Language Pre-training for Image Captioning
- MURAL: Multimodal, Multitask Retrieval Across Languages
- CCMB: A Large-scale Chinese Cross-modal Benchmark
- FooDI-ML: a large multi-language dataset of food, drinks and groceries images and descriptions
- Efficient large-scale image retrieval with deep feature orthogonality and Hybrid-Swin-Transformers
- "Wikily" Supervised Neural Translation Tailored to Cross-Lingual Tasks