2 papers
cs.CV2024
ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
Han Fang, Xianghao Zang, Chao Ban +5
Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since vid…
cs.CV2023
Mask to reconstruct: Cooperative Semantics Completion for Video-text Retrieval
Han Fang, Zhifei Yang, Xianghao Zang +2
Recently, masked video modeling has been widely explored and significantly improved the model's understanding ability of visual regions at a local level. However, existing methods…