Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
arXiv:1604.06838
Abstract
This paper strives to find the sentence best describing the content of an image or video. Different from existing works, which rely on a joint subspace for image / video to sentence matching, we propose to do so in a visual space only. We contribute Word2VisualVec, a deep neural network architecture that learns to predict a deep visual encoding of textual input based on sentence vectorization and a multi-layer perceptron. We thoroughly analyze its architectural design, by varying the sentence vectorization strategy, network depth and the deep feature to predict for image to sentence matching. We also generalize Word2VisualVec for matching a video to a sentence, by extending the predictive abilities to 3-D ConvNet features as well as a visual-audio representation. Experiments on four challenging image and video benchmarks detail Word2VisualVec's properties, capabilities for image and video to sentence matching, and on all datasets its state-of-the-art results.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Residual Learning for Image Recognition
- Show and Tell: A Neural Image Caption Generator
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Image Captioning with Deep Bidirectional LSTMs
- Jointly Modeling Embedding and Translation to Bridge Video and Language
Cited by in corpus (11)
- Use What You Have: Video Retrieval Using Representations From Collaborative Experts
- Audio Retrieval with Natural Language Queries: A Benchmark Study
- Action2Vec: A Crossmodal Embedding Approach to Action Learning
- A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
- Full-Network Embedding in a Multimodal Embedding Pipeline
- Memory Enhanced Embedding Learning for Cross-Modal Video-Text Retrieval
- Rethinking movie genre classification with fine-grained semantic clustering
- Activity Recognition on a Large Scale in Short Videos - Moments in Time Dataset
- Stacked Convolutional Deep Encoding Network for Video-Text Retrieval
- Interactive Video Retrieval with Dialog
- Image2song: Song Retrieval via Bridging Image Content and Lyric Words