1 paper
Ni Wang, Dongliang Liao, Xing Xu
Currently, in the field of video-text retrieval, there are many transformer-based methods. Most of them usually stack frame features and regrade frames as tokens, then use transfor…