2 papers
cs.IR2026
MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval
Chan Hur, SeungWoo Song, Jeong-hun Hong +3
Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they…
cs.CV2025
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
Chan Hur, Jeong-hun Hong, Dong-hun Lee +4
In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional…