4 citations · 5 across the 4 of their papers we have counts for
4 papers
Controllable Augmentations for Video Representation Learning
Rui Qian, Weiyao Lin, John See +1
This paper focuses on self-supervised video representation learning. Most existing approaches follow the contrastive learning pipeline to construct positive and negative pairs by s…
CLIP4Caption ++: Multi-CLIP for Video Caption
Mingkang Tang, Zhanyu Wang, Zhaoyang Zeng +2
This report describes our solution to the VALUE Challenge 2021 in the captioning task. Our solution, named CLIP4Caption++, is built on X-Linear/X-Transformer, which is an advanced…
CLIP4Caption: CLIP for Video Caption
Mingkang Tang, Zhanyu Wang, Zhenhua Liu +3
Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual r…
Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization
Rui Qian, Yuxi Li, Huabin Liu +5
The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics…