4 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 2 cited
Learning Commonsense-aware Moment-Text Alignment for Fast Video Temporal Grounding
Ziyue Wu, Junyu Gao, Shucheng Huang +1
Grounding temporal video segments described in natural language queries effectively and efficiently is a crucial capability needed in vision-and-language fields. In this paper, we…
cs.CV2022★ 4 cited
Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action Localization
Junyu Gao, Mengyuan Chen, Changsheng Xu
We target at the task of weakly-supervised action localization (WSAL), where only video-level action labels are available during model training. Despite the recent progress, existi…