64 citations · 162 across the 10 of their papers we have counts for
14 papers · 1 filter
Controllable Augmentations for Video Representation Learning
Rui Qian, Weiyao Lin, John See +1
This paper focuses on self-supervised video representation learning. Most existing approaches follow the contrastive learning pipeline to construct positive and negative pairs by s…
Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization
Rui Qian, Yuxi Li, Huabin Liu +5
The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics…
Variational Pedestrian Detection
Yuang Zhang, Huanyu He, Jianguo Li +3
Pedestrian detection in a crowd is a challenging task due to a high number of mutually-occluding human instances, which brings ambiguity and optimization difficulties to the curren…
Delving into the Cyclic Mechanism in Semi-supervised Video Object Segmentation
Yuxi Li, Ning Xu, Jinlong Peng +2
In this paper, we address several inadequacies of current video object segmentation pipelines. Firstly, a cyclic mechanism is incorporated to the standard semi-supervised process t…
Finding Action Tubes with a Sparse-to-Dense Framework
Yuxi Li, Weiyao Lin, Tao Wang +5
The task of spatial-temporal action detection has attracted increasing attention among researchers. Existing dominant methods solve this problem by relying on short-term informatio…
CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Yuxi Li, Weiyao Lin, John See +4
Most current pipelines for spatio-temporal action localization connect frame-wise or clip-wise detection results to generate action proposals, where only local information is explo…