1 paper
Pedro Rodriguez, Mahmoud Azab, Becka Silvert +4
Searching troves of videos with textual descriptions is a core multimodal retrieval task. Owing to the lack of a purpose-built dataset for text-to-video retrieval, video captioning…