5 papers · 1 filter
Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
Zhengqi Zhao, Xiaohu Huang, Hao Zhou +6
The key to action counting is accurately locating each video's repetitive actions. Instead of estimating the probability of each frame belonging to an action directly, we propose a…
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
Xiaohu Huang, Hao Zhou, Kun Yao +1
In this paper, we introduce FROSTER, an effective framework for open-vocabulary action recognition. The CLIP model has achieved remarkable success in a range of image-based tasks,…
HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception
Junkun Yuan, Xinyu Zhang, Hao Zhou +12
Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting…
Sign Language Translation with Iterative Prototype
Huijie Yao, Wengang Zhou, Hao Feng +3
This paper presents IP-SLT, a simple yet effective framework for sign language translation (SLT). Our IP-SLT adopts a recurrent structure and enhances the semantic representation (…
Towards Diverse Temporal Grounding under Single Positive Labels
Hao Zhou, Chongyang Zhang, Yanjun Chen +1
Temporal grounding aims to retrieve moments of the described event within an untrimmed video by a language query. Typically, existing methods assume annotations are precise and uni…