Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
arXiv:1411.2539
Abstract
Inspired by recent advances in multimodal learning and machine translation, we introduce an encoder-decoder pipeline that learns (a): a multimodal joint embedding space with images and text and (b): a novel language model for decoding distributed representations from our space. Our pipeline effectively unifies joint image-text embedding models with multimodal neural language models. We introduce the structure-content neural language model that disentangles the structure of a sentence to its content, conditioned on representations produced by the encoder. The encoder allows one to rank images and sentences while the decoder can generate novel descriptions from scratch. Using LSTM to encode sentences, we match the state-of-the-art performance on Flickr8K and Flickr30K without using object detections. We also set new best results when using the 19-layer Oxford convolutional network. Furthermore we show that with linear encoders, the learned embedding space captures multimodal regularities in terms of vector space arithmetic e.g. *image of a blue car* - "blue" + "red" is near images of red cars. Sample captions generated for 800 images are made available for comparison.
13 pages. NIPS 2014 deep learning workshop
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Recurrent Neural Network Regularization
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Explain Images with Multimodal Recurrent Neural Networks
- A Multiplicative Model for Learning Distributed Text-Based Attribute Representations
Cited by in corpus (21)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Generative Moment Matching Networks
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- Exploring Nearest Neighbor Approaches for Image Captioning
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- A Dataset for Movie Description
- Phrase-based Image Captioning
- Simple Image Description Generator via a Linear Phrase-Based Approach
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Generating Multi-Sentence Lingual Descriptions of Indoor Scenes
- Learning language through pictures
- Connectionist-Symbolic Machine Intelligence using Cellular Automata based Reservoir-Hyperdimensional Computing
- Learning Temporal Embeddings for Complex Video Analysis
- The Long-Short Story of Movie Description