Learning Deep Structure-Preserving Image-Text Embeddings
arXiv:1511.06078
Abstract
This paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a large margin objective that combines cross-view ranking constraints with within-view neighborhood structure preservation constraints inspired by metric learning literature. Extensive experiments show that our approach gains significant improvements in accuracy for image-to-text and text-to-image retrieval. Our method achieves new state-of-the-art results on the Flickr30K and MSCOCO image-sentence datasets and shows promise on the new task of phrase localization on the Flickr30K Entities dataset.
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- Computing the Stereo Matching Cost with a Convolutional Neural Network
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- Finding Linear Structure in Large Datasets with Scalable Canonical Correlation Analysis
Cited by in corpus (6)
- Multispectral Deep Neural Networks for Pedestrian Detection
- A Fast and Accurate One-Stage Approach to Visual Grounding
- HUSE: Hierarchical Universal Semantic Embeddings
- Multiple Visual-Semantic Embedding for Video Retrieval from Query Sentence
- ParNet: Position-aware Aggregated Relation Network for Image-Text matching
- T-EMDE: Sketching-based global similarity for cross-modal retrieval