PreViTS: Contrastive Pretraining with Video Tracking Supervision
arXiv:2112.00804
Abstract
Videos are a rich source for self-supervised learning (SSL) of visual representations due to the presence of natural temporal transformations of objects. However, current methods typically randomly sample video clips for learning, which results in an imperfect supervisory signal. In this work, we propose PreViTS, an SSL framework that utilizes an unsupervised tracking signal for selecting clips containing the same object, which helps better utilize temporal transformations of objects. PreViTS further uses the tracking signal to spatially constrain the frame regions to learn from and trains the model to locate meaningful objects by providing supervision on Grad-CAM attention maps. To evaluate our approach, we train a momentum contrastive (MoCo) encoder on VGG-Sound and Kinetics-400 datasets with PreViTS. Training with PreViTS outperforms representations learnt by contrastive strategy alone on video downstream tasks, obtaining state-of-the-art performance on action classification. PreViTS helps learn feature representations that are more robust to changes in background and context, as seen by experiments on datasets with background changes. Learning from large-scale videos with PreViTS could lead to more accurate and robust visual feature representations.
To be presented at WACV 2023
References in corpus (8)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Improved Baselines with Momentum Contrastive Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Self-supervised Co-training for Video Representation Learning
- Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition
- RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
- Object-aware Contrastive Learning for Debiased Scene Representation
- Unsupervised Object-Level Representation Learning from Scene Images