3 papers
cs.IR2026
MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval
Chan Hur, SeungWoo Song, Jeong-hun Hong +3
Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they…
cs.CV2026
G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation
Hojun Song, Chae-yeong Song, Jeong-hun Hong +5
Point cloud segmentation is critical for 3D scene understanding. However, sparse and irregular point distributions provide limited appearance evidence, making geometry-only feature…
cs.CV2025
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
Chan Hur, Jeong-hun Hong, Dong-hun Lee +4
In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional…