1 paper · 1 filter
Jiaxin Wu, Xiao-Yong Wei, Qing Li
The rise of short-form video platforms and the emergence of multimodal large language models (MLLMs) have amplified the need for scalable, effective, zero-shot text-to-video retrie…