2 papers
cs.CV2025
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
Zilin Du, Haoxin Li, Jianfei Yu +1
Visual grounding aims to localize the image regions based on a textual query. Given the difficulty of large-scale data curation, we investigate how to effectively learn visual grou…
cs.CL2024
Multilingual Synopses of Movie Narratives: A Dataset for Vision-Language Story Understanding
Yidan Sun, Jianfei Yu, Boyang Li
Story video-text alignment, a core task in computational story understanding, aims to align video clips with corresponding sentences in their descriptions. However, progress on the…