1 paper
Han Fang, Xianghao Zang, Chao Ban +5
Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since vid…