Commonly Uncommon: Semantic Sparsity in Situation Recognition
arXiv:1612.00901
Abstract
Semantic sparsity is a common challenge in structured visual classification problems; when the output space is complex, the vast majority of the possible predictions are rarely, if ever, seen in the training set. This paper studies semantic sparsity in situation recognition, the task of producing structured summaries of what is happening in images, including activities, objects and the roles objects play within the activity. For this problem, we find empirically that most object-role combinations are rare, and current state-of-the-art models significantly underperform in this sparse data regime. We avoid many such errors by (1) introducing a novel tensor composition function that learns to share examples across role-noun combinations and (2) semantically augmenting our training data with automatically gathered examples of rarely observed outputs using web data. When integrated within a complete CRF-based structured prediction model, the tensor-based approach outperforms existing state of the art by a relative improvement of 2.11% and 4.40% on top-5 verb and noun-role accuracy, respectively. Adding 5 million images with our semantic augmentation techniques gives further relative improvements of 6.23% and 9.57% on top-5 verb and noun-role accuracy.
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Explain Images with Multimodal Recurrent Neural Networks
- Visual Semantic Role Labeling
- Show and Tell: A Neural Image Caption Generator
- Learning a Recurrent Visual Representation for Image Caption Generation
- Exploring Nearest Neighbor Approaches for Image Captioning
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Visual Relationship Detection with Language Priors
- Learning to generalize to new compositions in image understanding
- Describing Common Human Visual Actions in Images