Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
arXiv:1505.04870
Abstract
The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. Such annotations are essential for continued progress in automatic image description and grounded language understanding. They enable us to define a new benchmark for localization of textual entity mentions in an image. We present a strong baseline for this task that combines an image-text embedding, detectors for common objects, a color classifier, and a bias towards selecting larger objects. While our baseline rivals in accuracy more complex state-of-the-art models, we show that its gains cannot be easily parlayed into improvements on such tasks as image-sentence retrieval, thus underlining the limitations of current methods and the need for further research.
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Learning a Recurrent Visual Representation for Image Caption Generation
- Language Models for Image Captioning: The Quirks and What Works
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- Multimodal Convolutional Neural Networks for Matching Image and Sentence
Cited by in corpus (21)
- ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
- Florence: A New Foundation Model for Computer Vision
- Order-Embeddings of Images and Language
- DenseCap: Fully Convolutional Localization Networks for Dense Captioning
- Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction
- Attention Correctness in Neural Image Captioning
- Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
- A Comprehensive Survey of Deep Learning for Image Captioning
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- Learning Cross-modal Context Graph for Visual Grounding
- Top-down Visual Saliency Guided by Captions
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- Neural Text Generation with Artificial Negative Examples
- Top-down Neural Attention by Excitation Backprop
- A Pipeline for Creative Visual Storytelling
- MULE: Multimodal Universal Language Embedding
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Visual Reasoning with Natural Language
- SPICE: Semantic Propositional Image Caption Evaluation
- From Route Instructions to Landmark Graphs
- Optimizing Open-Ended Crowdsourcing: The Next Frontier in Crowdsourced Data Management