4 citations · 9 across the 3 of their papers we have counts for
1 paper · 1 filter
Shen Yan, Xuehan Xiong, Arsha Nagrani +5
While large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos…