1 paper
Yicheng Duan, Xi Huang, Duo Chen
The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with ad…