Learning to generalize to new compositions in image understanding
arXiv:1608.07639
Abstract
Recurrent neural networks have recently been used for learning to describe images using natural language. However, it has been observed that these models generalize poorly to scenes that were not observed during training, possibly depending too strongly on the statistics of the text in the training data. Here we propose to describe images using short structured representations, aiming to capture the crux of a description. These structured representations allow us to tease-out and evaluate separately two types of generalization: standard generalization to new images with similar scenes, and generalization to new combinations of known entities. We compare two learning approaches on the MS-COCO dataset: a state-of-the-art recurrent network based on an LSTM (Show, Attend and Tell), and a simple structured prediction model on top of a deep network. We find that the structured model generalizes to new compositions substantially better than the LSTM, ~7 times the accuracy of predicting structured representations. By providing a concrete method to quantify generalization for unseen combinations, we argue that structured representations and compositional splits are a useful benchmark for image captioning, and advocate compositional models that capture linguistic and visual structure.
References in corpus (5)
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Explain Images with Multimodal Recurrent Neural Networks
- Learning a Recurrent Visual Representation for Image Caption Generation
- Latent Embeddings for Zero-shot Classification
Cited by in corpus (15)
- A Comprehensive Survey of Scene Graphs: Generation and Application
- Pixels to Graphs by Associative Embedding
- A causal view of compositional zero-shot recognition
- Visual Translation Embedding Network for Visual Relation Detection
- Context-Dependent Diffusion Network for Visual Relationship Detection
- On the Binding Problem in Artificial Neural Networks
- C-VQA: A Compositional Split of the Visual Question Answering (VQA) v1.0 Dataset
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Attentive Relational Networks for Mapping Images to Scene Graphs
- Adaptive Confidence Smoothing for Generalized Zero-Shot Learning
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- ViP-CNN: Visual Phrase Guided Convolutional Neural Network
- Learning Associative Inference Using Fast Weight Memory
- Learning of Colors from Color Names: Distribution and Point Estimation
- Shuffle-Then-Assemble: Learning Object-Agnostic Visual Relationship Features