Natural Language Object Retrieval
arXiv:1511.04164
Abstract
In this paper, we address the task of natural language object retrieval, to localize a target object within a given image based on a natural language query of the object. Natural language object retrieval differs from text-based image retrieval task as it involves spatial information about objects within the scene and global scene context. To address this issue, we propose a novel Spatial Context Recurrent ConvNet (SCRC) model as scoring function on candidate boxes for object retrieval, integrating spatial configurations and global scene-level contextual information into the network. Our model processes query text, local image descriptors, spatial configurations and global context features through a recurrent network, outputs the probability of the query text conditioned on each candidate box as a score for the box, and can transfer visual-linguistic knowledge from image captioning domain to our task. Experimental results demonstrate that our method effectively utilizes both local and global information, outperforming previous baseline methods significantly on different datasets and scenarios, and can exploit large scale vision and language datasets for knowledge transfer.
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Rich feature hierarchies for accurate object detection and semantic segmentation
- Show and Tell: A Neural Image Caption Generator
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Generation and Comprehension of Unambiguous Object Descriptions
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Segmentation from Natural Language Expressions
Cited by in corpus (15)
- Segmentation from Natural Language Expressions
- Real-Time Referring Expression Comprehension by Single-Stage Grounding Network
- VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation
- A Fast and Accurate One-Stage Approach to Visual Grounding
- Weakly-supervised Visual Grounding of Phrases with Linguistic Structures
- Generating Descriptions with Grounded and Co-Referenced People
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Grounding Referring Expressions in Images by Variational Context
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- Top-down Visual Saliency Guided by Captions
- Learning to Disambiguate by Asking Discriminative Questions
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions