See, Hear, and Read: Deep Aligned Representations
arXiv:1706.00932
Abstract
We capitalize on large amounts of readily-available, synchronous data to learn a deep discriminative representations shared across three major natural modalities: vision, sound and language. By leveraging over a year of sound from video and millions of sentences paired with images, we jointly train a deep convolutional network for aligned representation learning. Our experiments suggest that this representation is useful for several tasks, such as cross-modal retrieval or transferring classifiers between modalities. Moreover, although our network is only trained with image+text and image+sound pairs, it can transfer between text and sound as well, a transfer the network never observed during training. Visualizations of our representation reveal many hidden units which automatically emerge to detect concepts, independent of the modality.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Object Detectors Emerge in Deep Scene CNNs
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- SoundNet: Learning Sound Representations from Unlabeled Video
- Show and Tell: A Neural Image Caption Generator
Cited by in corpus (9)
- Not only Look, but also Listen: Learning Multimodal Violence Detection under Weak Supervision
- Self-supervised Moving Vehicle Tracking with Stereo Sound
- Deep Audio-Visual Learning: A Survey
- Deep Multimodal Feature Encoding for Video Ordering
- Self-supervised Audio Spatialization with Correspondence Classifier
- Unsupervised Generative Adversarial Alignment Representation for Sheet music, Audio and Lyrics
- Learning to Localize Sound Sources in Visual Scenes: Analysis and Applications
- Do We Need Sound for Sound Source Localization?
- Transcription-Enriched Joint Embeddings for Spoken Descriptions of Images and Videos