Modeling Context in Referring Expressions
arXiv:1608.00272
Abstract
Humans refer to objects in their environments all the time, especially in dialogue with other people. We explore generating and comprehending natural language referring expressions for objects in images. In particular, we focus on incorporating better measures of visual context into referring expression models and find that visual comparison to other objects within an image helps improve performance significantly. We also develop methods to tie the language generation process together, so that we generate expressions for all objects of a particular category jointly. Evaluation on three recent datasets - RefCOCO, RefCOCO+, and RefCOCOg, shows the advantages of our methods for both referring expression generation and comprehension.
19 pages, 6 figures, in ECCV 2016; authors, references and acknowledgement updated
References in corpus (2)
Cited by in corpus (12)
- ImageNet pre-trained models with batch normalization
- Temporal Context Network for Activity Localization in Videos
- VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- GuessWhat?! Visual object discovery through multi-modal dialogue
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Comprehension-guided referring expressions
- Grounding Spatio-Semantic Referring Expressions for Human-Robot Interaction
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Learning to Disambiguate by Asking Discriminative Questions
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries