Explain Images with Multimodal Recurrent Neural Networks
arXiv:1410.1090
Abstract
In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel sentence descriptions to explain the content of images. It directly models the probability distribution of generating a word given previous words and the image. Image descriptions are generated by sampling from this distribution. The model consists of two sub-networks: a deep recurrent neural network for sentences and a deep convolutional network for images. These two sub-networks interact with each other in a multimodal layer to form the whole m-RNN model. The effectiveness of our model is validated on three benchmark datasets (IAPR TC-12, Flickr 8K, and Flickr 30K). Our model outperforms the state-of-the-art generative method. In addition, the m-RNN model can be applied to retrieval tasks for retrieving images or sentences, and achieves significant performance improvement over the state-of-the-art methods which directly optimize the ranking objective function for retrieval.
References in corpus (1)
Cited by in corpus (30)
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- Learning a Recurrent Visual Representation for Image Caption Generation
- Exploring Nearest Neighbor Approaches for Image Captioning
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Learning to generalize to new compositions in image understanding
- Phrase-based Image Captioning
- Boosting Image Captioning with Attributes
- Identity-Aware Textual-Visual Matching with Latent Co-attention
- Simple Image Description Generator via a Linear Phrase-Based Approach
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Learning Semantic Concepts and Order for Image and Sentence Matching
- A Pooling Approach to Modelling Spatial Relations for Image Retrieval and Annotation
- Video Captioning with Transferred Semantic Attributes
- Generating Multi-Sentence Lingual Descriptions of Indoor Scenes
- A New Evaluation Protocol and Benchmarking Results for Extendable Cross-media Retrieval
- Multi-Modal Music Information Retrieval: Augmenting Audio-Analysis with Visual Computing for Improved Music Video Analysis
- Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive Style
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- DeepDiary: Automatic Caption Generation for Lifelogging Image Streams
- Transfer learning from language models to image caption generators: Better models may not transfer better
- Kite: Automatic speech recognition for unmanned aerial vehicles
- Improving Image Captioning by Leveraging Knowledge Graphs
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- AI-Powered Text Generation for Harmonious Human-Machine Interaction: Current State and Future Directions