128 citations · 240 across the 19 of their papers we have counts for
1 paper · 1 filter
Shuai Zhao, Linchao Zhu, Xiaohan Wang +1
Recently, large-scale pre-training methods like CLIP have made great progress in multi-modal research such as text-video retrieval. In CLIP, transformers are vital for modeling com…