Generation and Comprehension of Unambiguous Object Descriptions
arXiv:1511.02283
Abstract
We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods that generate descriptions of objects without taking into account other potentially ambiguous objects in the scene. Our model is inspired by recent successes of deep learning methods for image captioning, but while image captioning is difficult to evaluate, our task allows for easy objective evaluation. We also present a new large-scale dataset for referring expressions, based on MS-COCO. We have released the dataset and a toolbox for visualization and evaluation, see https://github.com/mjhucla/Google_Refexp_toolbox
We have released the Google Refexp dataset together with a toolbox for visualization and evaluation, see https://github.com/mjhucla/Google_Refexp_toolbox. Camera ready version for CVPR 2016
References in corpus (9)
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- LSTM: A Search Space Odyssey
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Exploring Nearest Neighbor Approaches for Image Captioning
- Robot Language Learning, Generation, and Comprehension
Cited by in corpus (21)
- ImageNet pre-trained models with batch normalization
- MAttNet: Modular Attention Network for Referring Expression Comprehension
- Multi-Modal Mutual Attention and Iterative Interaction for Referring Image Segmentation
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
- Natural Language Object Retrieval
- Real-Time Referring Expression Comprehension by Single-Stage Grounding Network
- Context-aware Captions from Context-agnostic Supervision
- A Fast and Accurate One-Stage Approach to Visual Grounding
- Weakly-supervised learning of visual relations
- Revisiting Image-Language Networks for Open-ended Phrase Detection
- Language Conditioned Spatial Relation Reasoning for 3D Object Grounding
- Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- Dense Captioning with Joint Inference and Visual Context
- Learning to Disambiguate by Asking Discriminative Questions
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Discriminative Bimodal Networks for Visual Localization and Detection with Natural Language Queries
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions