6 citations · 6 across the 1 of their papers we have counts for
1 paper
Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold +4
A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of s…