24 citations · 38 across the 4 of their papers we have counts for
5 papers · 1 filter
Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models
Dezhao Luo, Jiabo Huang, Shaogang Gong +2
Accurate video moment retrieval (VMR) requires universal visual-textual correlations that can handle unknown vocabulary and unseen scenes. However, the learned correlations are lik…
Video 3D Sampling for Self-supervised Representation Learning
Wei Li, Dezhao Luo, Bo Fang +2
Most of the existing video self-supervised methods mainly leverage temporal signals of videos, ignoring that the semantics of moving objects and environmental information are all c…
Exploring Relations in Untrimmed Videos for Self-Supervised Learning
Dezhao Luo, Bo Fang, Yu Zhou +3
Existing video self-supervised learning methods mainly rely on trimmed videos for model training. However, trimmed datasets are manually annotated from untrimmed videos. In this se…
Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning
Yuan Yao, Chang Liu, Dezhao Luo +2
In self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capa…
Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
Dezhao Luo, Chang Liu, Yu Zhou +4
We propose a novel self-supervised method, referred to as Video Cloze Procedure (VCP), to learn rich spatial-temporal representations. VCP first generates "blanks" by withholding v…