Look, Listen and Learn
arXiv:1705.08168
Abstract
We consider the question: what can be learnt by looking at and listening to a large number of unlabelled videos? There is a valuable, but so far untapped, source of information contained in the video itself -- the correspondence between the visual and the audio streams, and we introduce a novel "Audio-Visual Correspondence" learning task that makes use of this. Training visual and audio networks from scratch, without any additional supervision other than the raw unconstrained videos themselves, is shown to successfully solve this task, and, more interestingly, result in good visual and audio representations. These features set the new state-of-the-art on two sound classification benchmarks, and perform on par with the state-of-the-art self-supervised approaches on ImageNet classification. We also demonstrate that the network is able to localize objects in both modalities, as well as perform fine-grained recognition tasks.
Appears in: IEEE International Conference on Computer Vision (ICCV) 2017
References in corpus (4)
Cited by in corpus (19)
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Towards Learning a Universal Non-Semantic Representation of Speech
- Transfer learning for music classification and regression tasks
- Video Understanding as Machine Translation
- Nonlinear Invariant Risk Minimization: A Causal Approach
- Self-supervised Moving Vehicle Tracking with Stereo Sound
- Emotion Recognition in Speech using Cross-Modal Transfer in the Wild
- Mining YouTube - A dataset for learning fine-grained action concepts from webly supervised video data
- Listen to Look: Action Recognition by Previewing Audio
- Fast forwarding Egocentric Videos by Listening and Watching
- Unsupervised Learning of Semantic Audio Representations
- Adaptive pooling operators for weakly labeled sound event detection
- Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows
- Assessing the Contribution of Semantic Congruency to Multisensory Integration and Conflict Resolution
- Towards Robust Pattern Recognition: A Review
- Rethinking the constraints of multimodal fusion: case study in Weakly-Supervised Audio-Visual Video Parsing
- Translating Paintings Into Music Using Neural Networks
- Substitute Teacher Networks: Learning with Almost No Supervision
- 3D-MOV: Audio-Visual LSTM Autoencoder for 3D Reconstruction of Multiple Objects from Video