35 citations · 41 across the 3 of their papers we have counts for
4 papers · 1 filter
Masked Contrastive Pre-Training for Efficient Video-Text Retrieval
Fangxun Shu, Biaolong Chen, Yue Liao +6
We present a simple yet effective end-to-end Video-language Pre-training (VidLP) framework, Masked Contrastive Video-language Pretraining (MAC), for video-text retrieval tasks. Our…
Mining the Benefits of Two-stage and One-stage HOI Detection
Aixi Zhang, Yue Liao, Si Liu +4
Two-stage methods have dominated Human-Object Interaction (HOI) detection for several years. Recently, one-stage HOI detection methods have become popular. In this paper, we aim to…
Multi-Granularity Network with Modal Attention for Dense Affective Understanding
Baoming Yan, Lin Wang, Ke Gao +5
Video affective understanding, which aims to predict the evoked expressions by the video content, is desired for video creation and recommendation. In the recent EEV challenge, a d…
Augmented Bi-path Network for Few-shot Learning
Baoming Yan, Chen Zhou, Bo Zhao +5
Few-shot Learning (FSL) which aims to learn from few labeled training data is becoming a popular research topic, due to the expensive labeling cost in many real-world applications.…