Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
arXiv:1411.2539
Abstract
Inspired by recent advances in multimodal learning and machine translation, we introduce an encoder-decoder pipeline that learns (a): a multimodal joint embedding space with images and text and (b): a novel language model for decoding distributed representations from our space. Our pipeline effectively unifies joint image-text embedding models with multimodal neural language models. We introduce the structure-content neural language model that disentangles the structure of a sentence to its content, conditioned on representations produced by the encoder. The encoder allows one to rank images and sentences while the decoder can generate novel descriptions from scratch. Using LSTM to encode sentences, we match the state-of-the-art performance on Flickr8K and Flickr30K without using object detections. We also set new best results when using the 19-layer Oxford convolutional network. Furthermore we show that with linear encoders, the learned embedding space captures multimodal regularities in terms of vector space arithmetic e.g. *image of a blue car* - "blue" + "red" is near images of red cars. Sample captions generated for 800 images are made available for comparison.
13 pages. NIPS 2014 deep learning workshop
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Recurrent Neural Network Regularization
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Explain Images with Multimodal Recurrent Neural Networks
- A Multiplicative Model for Learning Distributed Text-Based Attribute Representations
Cited by in corpus (132)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Learning Transferable Visual Models From Natural Language Supervision
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Generative Moment Matching Networks
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- Detecting Sarcasm in Multimodal Social Platforms
- A Comprehensive Survey on Cross-modal Retrieval
- An Actor-Critic Algorithm for Sequence Prediction
- Learning What and Where to Draw
- Exploring Nearest Neighbor Approaches for Image Captioning
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Root Mean Square Layer Normalization
- Fluency-Guided Cross-Lingual Image Captioning
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- See, Hear, and Read: Deep Aligned Representations
- Cross-Modal Music Retrieval and Applications: An Overview of Key Methodologies
- Learning to generalize to new compositions in image understanding
- Modeling Context in Referring Expressions
- A Dataset for Movie Description
- Phrase-based Image Captioning
- Dual Attention Networks for Multimodal Reasoning and Matching
- DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis
- Action2Vec: A Crossmodal Embedding Approach to Action Learning
- Automatic Spatially-aware Fashion Concept Discovery
- 3G structure for image caption generation
- DeepStory: Video Story QA by Deep Embedded Memory Networks
- Teaching Machines to Describe Images via Natural Language Feedback
- Multi-modal gated recurrent units for image description
- TediGAN: Text-Guided Diverse Face Image Generation and Manipulation
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Simple Image Description Generator via a Linear Phrase-Based Approach
- Incorporating Global Visual Features into Attention-Based Neural Machine Translation
- M-VAD Names: a Dataset for Video Captioning with Naming
- Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning
- Learning Semantic Concepts and Order for Image and Sentence Matching
- Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
- A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
- Semantic Compositional Networks for Visual Captioning
- Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching
- Learning Audio - Sheet Music Correspondences for Score Identification and Offline Alignment
- Position Focused Attention Network for Image-Text Matching
- Transitive Hashing Network for Heterogeneous Multimedia Retrieval
- Multilingual Multi-modal Embeddings for Natural Language Processing
- Generating Multi-Sentence Lingual Descriptions of Indoor Scenes
- Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval
- Open-World Visual Recognition Using Knowledge Graphs
- A New Evaluation Protocol and Benchmarking Results for Extendable Cross-media Retrieval
- Cross-modal Scene Graph Matching for Relationship-aware Image-Text Retrieval
- Deep Bayesian Active Learning for Multiple Correct Outputs
- Step-Wise Hierarchical Alignment Network for Image-Text Matching
- Connectionist-Symbolic Machine Intelligence using Cellular Automata based Reservoir-Hyperdimensional Computing
- Learning language through pictures
- HANet: Hierarchical Alignment Networks for Video-Text Retrieval
- Understanding Image and Text Simultaneously: a Dual Vision-Language Machine Comprehension Task
- Better Text Understanding Through Image-To-Text Transfer
- LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval
- Improving Visually Grounded Sentence Representations with Self-Attention
- Tensor Product Generation Networks for Deep NLP Modeling
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- Text-guided Attention Model for Image Captioning
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
- Informative Image Captioning with External Sources of Information
- Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models
- Full-Network Embedding in a Multimodal Embedding Pipeline
- Weakly supervised cross-domain alignment with optimal transport
- Exploring Uncertainty Measures for Image-Caption Embedding-and-Retrieval Task
- Joint Visual-Textual Embedding for Multimodal Style Search
- Neural Storyboard Artist: Visualizing Stories with Coherent Image Sequences
- Self-Guiding Multimodal LSTM - when we do not have a perfect training dataset for image captioning
- Gaussian Smoothen Semantic Features (GSSF) -- Exploring the Linguistic Aspects of Visual Captioning in Indian Languages (Bengali) Using MSCOCO Framework
- Phrase-based Image Captioning with Hierarchical LSTM Model
- Learning Soft-Attention Models for Tempo-invariant Audio-Sheet Music Retrieval
- Improving Sales Forecasting Accuracy: A Tensor Factorization Approach with Demand Awareness
- UniVSE: Robust Visual Semantic Embeddings via Structured Semantic Representations
- Image Captioning using Deep Stacked LSTMs, Contextual Word Embeddings and Data Augmentation
- HUSE: Hierarchical Universal Semantic Embeddings
- Ladder Loss for Coherent Visual-Semantic Embedding
- Learning Temporal Embeddings for Complex Video Analysis
- MHSAN: Multi-Head Self-Attention Network for Visual Semantic Embedding
- Rethinking movie genre classification with fine-grained semantic clustering
- Personalized Multimodal Feedback Generation in Education
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Evaluating Text-to-Image Matching using Binary Image Selection (BISON)
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval
- 3D Human motion anticipation and classification
- Upgrading the Newsroom: An Automated Image Selection System for News Articles
- What is Learned in Visually Grounded Neural Syntax Acquisition
- An Unsupervised Sampling Approach for Image-Sentence Matching Using Document-Level Structural Information
- Rudder: A Cross Lingual Video and Text Retrieval Dataset
- A Universal Model for Cross Modality Mapping by Relational Reasoning
- Exploiting Visual Semantic Reasoning for Video-Text Retrieval
- Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations
- A Strong and Robust Baseline for Text-Image Matching
- Target-Oriented Deformation of Visual-Semantic Embedding Space
- Beyond the Deep Metric Learning: Enhance the Cross-Modal Matching with Adversarial Discriminative Domain Regularization
- SSAN: Separable Self-Attention Network for Video Representation Learning
- Learning Relation Alignment for Calibrated Cross-modal Retrieval
- ActBERT: Learning Global-Local Video-Text Representations
- TVDIM: Enhancing Image Self-Supervised Pretraining via Noisy Text Data
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- Big Data driven Product Design: A Survey
- Expressing Objects just like Words: Recurrent Visual Embedding for Image-Text Matching
- Automated Knee X-ray Report Generation
- Exploring the Challenges towards Lifelong Fact Learning
- OptiBox: Breaking the Limits of Proposals for Visual Grounding
- The Long-Short Story of Movie Description
- Deep Unified Multimodal Embeddings for Understanding both Content and Users in Social Media Networks
- Retrieving and Highlighting Action with Spatiotemporal Reference
- YouMakeup VQA Challenge: Towards Fine-grained Action Understanding in Domain-Specific Videos
- Image2song: Song Retrieval via Bridging Image Content and Lyric Words
- Watch and Learn: Mapping Language and Noisy Real-world Videos with Self-supervision
- Contrastive Learning of Visual-Semantic Embeddings
- Improving Zero-shot Multilingual Neural Machine Translation for Low-Resource Languages
- An animated picture says at least a thousand words: Selecting Gif-based Replies in Multimodal Dialog
- The State of the Art when using GPUs in Devising Image Generation Methods Using Deep Learning
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- Attribute Guided Sparse Tensor-Based Model for Person Re-Identification
- PhotoChat: A Human-Human Dialogue Dataset with Photo Sharing Behavior for Joint Image-Text Modeling
- Case Relation Transformer: A Crossmodal Language Generation Model for Fetching Instructions
- PUNCH: Positive UNlabelled Classification based information retrieval in Hyperspectral images