550 citations · 627 across the 7 of their papers we have counts for
1 paper · 1 filter
Yili Li, Gang Xiong, Gaopeng Gou +4
Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrat…