Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
arXiv:1406.5679
Abstract
We introduce a model for bidirectional retrieval of images and sentences through a multi-modal embedding of visual and natural language data. Unlike previous models that directly map images or sentences into a common embedding space, our model works on a finer level and embeds fragments of images (objects) and fragments of sentences (typed dependency tree relations) into a common space. In addition to a ranking objective seen in previous work, this allows us to add a new fragment alignment objective that learns to directly associate these fragments across modalities. Extensive experimental evaluation shows that reasoning on both the global level of images and sentences and the finer level of their respective fragments significantly improves performance on image-sentence retrieval tasks. Additionally, our model provides interpretable predictions since the inferred inter-modal fragment alignment is explicit.
References in corpus (1)
Cited by in corpus (30)
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Exploring Nearest Neighbor Approaches for Image Captioning
- Recurrent Topic-Transition GAN for Visual Paragraph Generation
- Interleaved Text/Image Deep Mining on a Large-Scale Radiology Database for Automated Image Interpretation
- Dual Attention Networks for Multimodal Reasoning and Matching
- Visual Question Answering: A Survey of Methods and Datasets
- Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
- Multi-modal gated recurrent units for image description
- Multimodal Transformer with Multi-View Visual Representation for Image Captioning
- A Pooling Approach to Modelling Spatial Relations for Image Retrieval and Annotation
- Predicting Aircraft Trajectories: A Deep Generative Convolutional Recurrent Neural Networks Approach
- Hierarchical LSTMs with Adaptive Attention for Visual Captioning
- A New Evaluation Protocol and Benchmarking Results for Extendable Cross-media Retrieval
- Learning Structured Semantic Embeddings for Visual Recognition
- Deep Binaries: Encoding Semantic-Rich Cues for Efficient Textual-Visual Cross Retrieval
- Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning
- Full-Network Embedding in a Multimodal Embedding Pipeline
- Discover and Learn New Objects from Documentaries
- Joint Visual-Textual Embedding for Multimodal Style Search
- Rethinking the Artificial Neural Networks: A Mesh of Subnets with a Central Mechanism for Storing and Predicting the Data
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Hierarchical Photo-Scene Encoder for Album Storytelling
- Multitask Text-to-Visual Embedding with Titles and Clickthrough Data
- Improving Image Captioning by Leveraging Knowledge Graphs
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Self-view Grounding Given a Narrated 360° Video
- A Weighted Multi-Criteria Decision Making Approach for Image Captioning
- Joint Learning of Distributed Representations for Images and Texts