1 paper
Leqi Shen, Tianxiang Hao, Tao He +5
Most text-video retrieval methods utilize the text-image pre-trained models like CLIP as a backbone. These methods process each sampled frame independently by the image encoder, re…