Motion-aware Contrastive Video Representation Learning via Foreground-background Merging
arXiv:2109.15130
Abstract
In light of the success of contrastive learning in the image domain, current self-supervised video representation learning methods usually employ contrastive loss to facilitate video representation learning. When naively pulling two augmented views of a video closer, the model however tends to learn the common static background as a shortcut but fails to capture the motion information, a phenomenon dubbed as background bias. Such bias makes the model suffer from weak generalization ability, leading to worse performance on downstream tasks such as action recognition. To alleviate such bias, we propose \textbf{F}oreground-b\textbf{a}ckground \textbf{Me}rging (FAME) to deliberately compose the moving foreground region of the selected video onto the static background of others. Specifically, without any off-the-shelf detector, we extract the moving foreground out of background regions via the frame difference and color statistics, and shuffle the background regions among the videos. By leveraging the semantic consistency between the original clips and the fused ones, the model focuses more on the motion patterns and is debiased from the background shortcut. Extensive experiments demonstrate that FAME can effectively resist background cheating and thus achieve the state-of-the-art performance on downstream tasks across UCF101, HMDB51, and Diving48 datasets. The code and configurations are released at https://github.com/Mark12Ding/FAME.
CVPR2022 camera ready
References in corpus (14)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Bootstrap your own latent: A new approach to self-supervised Learning
- Improved Baselines with Momentum Contrastive Learning
- Exploring Simple Siamese Representation Learning
- Self-supervised Co-training for Video Representation Learning
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition
- Spatiotemporal Contrastive Video Representation Learning
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Self-supervised Video Representation Learning by Pace Prediction
- MixCo: Mix-up Contrastive Learning for Visual Representation
- Cycle-Contrast for Self-Supervised Video Representation Learning
- RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
- Motion-Focused Contrastive Learning of Video Representations