Multimodal Convolutional Neural Networks for Matching Image and Sentence
arXiv:1504.06063
Abstract
In this paper, we propose multimodal convolutional neural networks (m-CNNs) for matching image and sentence. Our m-CNN provides an end-to-end framework with convolutional architectures to exploit image representation, word composition, and the matching relations between the two modalities. More specifically, it consists of one image CNN encoding the image content, and one matching CNN learning the joint representation of image and sentence. The matching CNN composes words to different semantic fragments and learns the inter-modal relations between image and the composed fragments at different levels, thus fully exploit the matching relations between image and sentence. Experimental results on benchmark databases of bidirectional image and sentence retrieval demonstrate that the proposed m-CNNs can effectively capture the information necessary for image and sentence matching. Specifically, our proposed m-CNNs for bidirectional image and sentence retrieval on Flickr30K and Microsoft COCO databases achieve the state-of-the-art performances.
Accepted by ICCV 2015
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Improving neural networks by preventing co-adaptation of feature detectors
- Natural Language Processing (almost) from Scratch
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
- Going Deeper with Convolutions
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Explain Images with Multimodal Recurrent Neural Networks
- Show and Tell: A Neural Image Caption Generator
- Learning a Recurrent Visual Representation for Image Caption Generation
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
Cited by in corpus (33)
- Neural Models for Information Retrieval
- CNN+CNN: Convolutional Decoders for Image Captioning
- Discriminability objective for training descriptive captions
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- Emergent Translation in Multi-Agent Communication
- Dual Attention Networks for Multimodal Reasoning and Matching
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- Neural Information Retrieval: A Literature Review
- Reconstruction Network for Video Captioning
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- Learning to Answer Questions From Image Using Convolutional Neural Network
- Learning Semantic Concepts and Order for Image and Sentence Matching
- Towards Open-World Text-Guided Face Image Generation and Manipulation
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching
- Consensus-Aware Visual-Semantic Embedding for Image-Text Matching
- Comprehension-guided referring expressions
- AMC: Attention guided Multi-modal Correlation Learning for Image Search
- Learning semantic sentence representations from visually grounded language without lexical knowledge
- An In-Depth Analysis of Visual Tracking with Siamese Neural Networks
- Step-Wise Hierarchical Alignment Network for Image-Text Matching
- Instance-aware Image and Sentence Matching with Selective Multimodal LSTM
- Gated Hierarchical Attention for Image Captioning
- Deep Binaries: Encoding Semantic-Rich Cues for Efficient Textual-Visual Cross Retrieval
- Full-Network Embedding in a Multimodal Embedding Pipeline
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- Order embeddings and character-level convolutions for multimodal alignment
- Luandri: a Clean Lua Interface to the Indri Search Engine
- Leveraging Visual Question Answering for Image-Caption Ranking
- Holistic Multi-modal Memory Network for Movie Question Answering
- Exploration on Grounded Word Embedding: Matching Words and Images with Image-Enhanced Skip-Gram Model
- Listening while Speaking and Visualizing: Improving ASR through Multimodal Chain
- ParNet: Position-aware Aggregated Relation Network for Image-Text matching