1 paper
Chan Hur, Jeong-hun Hong, Dong-hun Lee +4
In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional…