Image Captioning with Deep Bidirectional LSTMs
arXiv:1604.00790
Abstract
This work presents an end-to-end trainable deep bidirectional LSTM (Long-Short Term Memory) model for image captioning. Our model builds on a deep convolutional neural network (CNN) and two separate LSTM networks. It is capable of learning long term visual-language interactions by making use of history and future context information at high level semantic space. Two novel deep bidirectional variant models, in which we increase the depth of nonlinearity transition in different way, are proposed to learn hierarchical visual-language embeddings. Data augmentation techniques such as multi-crop, multi-scale and vertical mirror are proposed to prevent overfitting in training deep models. We visualize the evolution of bidirectional LSTM internal states over time and qualitatively analyze how our models "translate" image to sentence. Our proposed models are evaluated on caption generation and image-sentence retrieval tasks with three benchmark datasets: Flickr8K, Flickr30K and MSCOCO datasets. We demonstrate that bidirectional LSTM models achieve highly competitive performance to the state-of-the-art results on caption generation even without integrating additional mechanism (e.g. object detection, attention model etc.) and significantly outperform recent methods on retrieval task.
accepted by ACMMM 2016 as full paper and oral presentation
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
Cited by in corpus (7)
- Representation Learning for Natural Language Processing
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- A Comprehensive Survey of Deep Learning for Image Captioning
- Full-Network Embedding in a Multimodal Embedding Pipeline
- Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-in-the-Blank Image Captioning
- Netizen-Style Commenting on Fashion Photos: Dataset and Diversity Measures
- Feature Fusion Effects of Tensor Product Representation on (De)Compositional Network for Caption Generation for Images