1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Xiangpeng Yang, Linchao Zhu, Xiaohan Wang +1
Text-video retrieval is a critical multi-modal task to find the most relevant video for a text query. Although pretrained models like CLIP have demonstrated impressive potential in…