Recurrent Models for Situation Recognition
arXiv:1703.06233
Abstract
This work proposes Recurrent Neural Network (RNN) models to predict structured 'image situations' -- actions and noun entities fulfilling semantic roles related to the action. In contrast to prior work relying on Conditional Random Fields (CRFs), we use a specialized action prediction network followed by an RNN for noun prediction. Our system obtains state-of-the-art accuracy on the challenging recent imSitu dataset, beating CRF-based models, including ones trained with additional data. Further, we show that specialized features learned from situation prediction can be transferred to the task of image captioning to more accurately describe human-object interactions.
To appear at ICCV 2017
References in corpus (15)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Recurrent Neural Network Regularization
- VQA: Visual Question Answering
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Image Captioning with Semantic Attention
- Show and Tell: A Neural Image Caption Generator
- CIDEr: Consensus-based Image Description Evaluation
- Boosting Image Captioning with Attributes
- What value do explicit high level concepts have in vision to language problems?
- Describing Common Human Visual Actions in Images
- Fine-grained Activity Recognition with Holistic and Pose based Features
- Commonly Uncommon: Semantic Sparsity in Situation Recognition