TCLR: Temporal Contrastive Learning for Video Representation
arXiv:2101.07974 · doi:10.1016/j.cviu.2022.103406
Abstract
Contrastive learning has nearly closed the gap between supervised and self-supervised learning of image representations, and has also been explored for videos. However, prior work on contrastive learning for video data has not explored the effect of explicitly encouraging the features to be distinct across the temporal dimension. We develop a new temporal contrastive learning framework consisting of two novel losses to improve upon existing contrastive self-supervised video representation learning methods. The local-local temporal contrastive loss adds the task of discriminating between non-overlapping clips from the same video, whereas the global-local temporal contrastive aims to discriminate between timesteps of the feature map of an input clip in order to increase the temporal diversity of the learned features. Our proposed temporal contrastive learning framework achieves significant improvement over the state-of-the-art results in various downstream video understanding tasks such as action recognition, limited-label action classification, and nearest-neighbor video retrieval on multiple video datasets and backbones. We also demonstrate significant improvement in fine-grained action classification for visually similar classes. With the commonly used 3D ResNet-18 architecture with UCF101 pretraining, we achieve 82.4\% (+5.1\% increase over the previous best) top-1 accuracy on UCF101 and 52.9\% (+5.4\% increase) on HMDB51 action classification, and 56.2\% (+11.7\% increase) Top-1 Recall on UCF101 nearest neighbor video retrieval. Code released at github.com/DAVEISHAN/TCLR.
Accepted to Computer Vision and Image Understanding (CVIU) Journal
References in corpus (22)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Bootstrap your own latent: A new approach to self-supervised Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Learning Representations by Maximizing Mutual Information Across Views
- Self-supervised Co-training for Video Representation Learning
- Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework
- Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition
- Would Mega-scale Datasets Further Enhance Spatiotemporal 3D CNNs?
- Spatiotemporal Contrastive Video Representation Learning
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Self-supervised Video Representation Learning by Pace Prediction
- Video Representation Learning with Visual Tempo Consistency
- Can Temporal Information Help with Contrastive Self-Supervised Learning?
- Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics
- VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
- Unsupervised Learning of Video Representations via Dense Trajectory Clustering
- RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
- Representation Learning with Video Deep InfoMax
- VIPriors 1: Visual Inductive Priors for Data-Efficient Deep Learning Challenges
- "Knights": First Place Submission for VIPriors21 Action Recognition Challenge at ICCV 2021
Cited by in corpus (11)
- Self-Supervised Learning for Videos: A Survey
- Contrastive encoder pre-training-based clustered federated learning for heterogeneous data
- Similarity Contrastive Estimation for Image and Video Soft Contrastive Self-Supervised Learning
- Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
- SudokuSens: Enhancing Deep Learning Robustness for IoT Sensing Applications using a Generative Approach
- Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
- Do We Really Need to Learn Representations from In-domain Data for Outlier Detection?
- Self-supervised Video-centralised Transformer for Video Face Clustering
- "Knights": First Place Submission for VIPriors21 Action Recognition Challenge at ICCV 2021
- Advancing Video Self-Supervised Learning via Image Foundation Models
- Overcoming the Domain Gap in Contrastive Learning of Neural Action Representations