UBoCo : Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection
arXiv:2111.14799
Abstract
Generic Event Boundary Detection (GEBD) is a newly suggested video understanding task that aims to find one level deeper semantic boundaries of events. Bridging the gap between natural human perception and video understanding, it has various potential applications, including interpretable and semantically valid video parsing. Still at an early development stage, existing GEBD solvers are simple extensions of relevant video understanding tasks, disregarding GEBD's distinctive characteristics. In this paper, we propose a novel framework for unsupervised/supervised GEBD, by using the Temporal Self-similarity Matrix (TSM) as the video representation. The new Recursive TSM Parsing (RTP) algorithm exploits local diagonal patterns in TSM to detect boundaries, and it is combined with the Boundary Contrastive (BoCo) loss to train our encoder to generate more informative TSMs. Our framework can be applied to both unsupervised and supervised settings, with both achieving state-of-the-art performance by a huge margin in GEBD benchmark. Especially, our unsupervised method outperforms the previous state-of-the-art "supervised" model, implying its exceptional efficacy.
References in corpus (9)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Bootstrap your own latent: A new approach to self-supervised Learning
- The Kinetics Human Action Video Dataset
- Improved Baselines with Momentum Contrastive Learning
- MLP-Mixer: An all-MLP Architecture for Vision
- Big Self-Supervised Models are Strong Semi-Supervised Learners
- Supervised Contrastive Learning
- Shot Contrastive Self-Supervised Learning for Scene Boundary Detection
- Generic Event Boundary Detection: A Benchmark for Event Segmentation