Visual Madlibs: Fill in the blank Image Generation and Question Answering
arXiv:1506.00278
Abstract
In this paper, we introduce a new dataset consisting of 360,001 focused natural language descriptions for 10,738 images. This dataset, the Visual Madlibs dataset, is collected using automatically produced fill-in-the-blank templates designed to gather targeted descriptions about: people and objects, their appearances, activities, and interactions, as well as inferences about the general scene or its broader context. We provide several analyses of the Visual Madlibs dataset and demonstrate its applicability to two new description generation tasks: focused description generation, and multiple-choice question-answering for images. Experiments using joint-embedding and deep learning methods show promising results on these tasks.
10 pages; 8 figures; 4 tables
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Natural Language Processing (almost) from Scratch
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Learning a Recurrent Visual Representation for Image Caption Generation
- Question Answering with Subgraph Embeddings
- Open Question Answering with Weakly Supervised Embedding Models
Cited by in corpus (9)
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Visual Question Answering: A Survey of Methods and Datasets
- Video Fill in the Blank with Merging LSTMs
- VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
- Visual Question: Predicting If a Crowd Will Agree on the Answer
- TAB-VCR: Tags and Attributes based Visual Commonsense Reasoning Baselines
- Mean Box Pooling: A Rich Image Representation and Output Embedding for the Visual Madlibs Task
- Visual Question Answering as a Multi-Task Problem
- Reasoning Over History: Context Aware Visual Dialog