3 papers
cs.CV2025
VoCap: Video Object Captioning and Segmentation from Any Prompt
Jasper Uijlings, Xingyi Zhou, Xiuye Gu +5
Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose…
cs.LG2025
Neptune: The Long Orbit to Benchmarking Long Video Understanding
Arsha Nagrani, Mingda Zhang, Ramin Mehran +10
We introduce Neptune, a benchmark for long video understanding that requires reasoning over long time horizons and across different modalities. Many existing video datasets and mod…
cs.CV2024
Visual Lexicon: Rich Image Features in Language Space
XuDong Wang, Xingyi Zhou, Alireza Fathi +2
We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intricate visual details that are of…