Self-supervised Learning for Video Correspondence Flow
arXiv:1905.00875
Abstract
The objective of this paper is self-supervised learning of feature embeddings that are suitable for matching correspondences along the videos, which we term correspondence flow. By leveraging the natural spatial-temporal coherence in videos, we propose to train a ``pointer'' that reconstructs a target frame by copying pixels from a reference frame. We make the following contributions: First, we introduce a simple information bottleneck that forces the model to learn robust features for correspondence matching, and prevent it from learning trivial solutions, \eg matching based on low-level colour information. Second, to tackle the challenges from tracker drifting, due to complex object deformations, illumination changes and occlusions, we propose to train a recursive model over long temporal windows with scheduled sampling and cycle consistency. Third, we achieve state-of-the-art performance on DAVIS 2017 video segmentation and JHMDB keypoint tracking tasks, outperforming all previous self-supervised learning approaches by a significant margin. Fourth, in order to shed light on the potential of self-supervised learning on the task of video correspondence flow, we probe the upper bound by training on additional data, \ie more diverse videos, further demonstrating significant improvements on video segmentation.
BMVC2019 (Oral Presentation)
References in corpus (3)
Cited by in corpus (23)
- Self-supervised Co-training for Video Representation Learning
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Space-Time Correspondence as a Contrastive Random Walk
- Video Representation Learning by Dense Predictive Coding
- Learning Video Object Segmentation from Unlabeled Videos
- Learning Video Representations from Textual Web Supervision
- CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions
- Probing the State of the Art: A Critical Look at Visual Representation Evaluation
- MAST: A Memory-Augmented Self-supervised Tracker
- Betrayed by Motion: Camouflaged Object Discovery via Motion Segmentation
- Correspondence Networks with Adaptive Neighbourhood Consensus
- Self-supervised Video Object Segmentation by Motion Grouping
- RANSAC-Flow: generic two-stage image alignment
- Breaking Shortcut: Exploring Fully Convolutional Cycle-Consistency for Video Correspondence Learning
- Group Collaborative Learning for Co-Salient Object Detection
- HDRVideo-GAN: Deep Generative HDR Video Reconstruction
- Self-supervised Object Tracking with Cycle-consistent Siamese Networks
- Progressive Temporal Feature Alignment Network for Video Inpainting
- Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- Sli2Vol: Annotate a 3D Volume from a Single Slice with Self-Supervised Learning
- Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
- DyStaB: Unsupervised Object Segmentation via Dynamic-Static Bootstrapping