1 paper · 1 filter
Minghan Li, Tongna Chen, Tianrui Lv +3
Existing text-to-video retrieval benchmarks are dominated by real-world footage where much of the semantics can be inferred from a single frame, leaving temporal reasoning and expl…