Learning Two-Branch Neural Networks for Image-Text Matching Tasks
arXiv:1704.03470
Abstract
Image-language matching tasks have recently attracted a lot of attention in the computer vision field. These tasks include image-sentence matching, i.e., given an image query, retrieving relevant sentences and vice versa, and region-phrase matching or visual grounding, i.e., matching a phrase to relevant regions. This paper investigates two-branch neural networks for learning the similarity between these two data modalities. We propose two network structures that produce different output representations. The first one, referred to as an embedding network, learns an explicit shared latent embedding space with a maximum-margin ranking loss and novel neighborhood constraints. Compared to standard triplet sampling, we perform improved neighborhood sampling that takes neighborhood information into consideration while constructing mini-batches. The second network structure, referred to as a similarity network, fuses the two branches via element-wise product and is trained with regression loss to directly predict a similarity score. Extensive experiments show that our networks achieve high accuracies for phrase localization on the Flickr30K Entities dataset and for bi-directional image-sentence retrieval on Flickr30K and MSCOCO datasets.
accepted version in TPAMI 2018
References in corpus (11)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distributed Representations of Words and Phrases and their Compositionality
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- Computing the Stereo Matching Cost with a Convolutional Neural Network
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Revisiting Batch Normalization For Practical Domain Adaptation
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Multimodal Convolutional Neural Networks for Matching Image and Sentence
Cited by in corpus (7)
- MAttNet: Modular Attention Network for Referring Expression Comprehension
- Discriminability objective for training descriptive captions
- Revisiting Image-Language Networks for Open-ended Phrase Detection
- Deep Matching Autoencoders
- Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint
- Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking
- Learning Shared Semantic Space with Correlation Alignment for Cross-modal Event Retrieval