papers
Publications (3)
cs.CV2025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Ye Wang, Ziheng Wang, Boshen Xu +14
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vis…
cs.CV2023
Beyond Domain Gap: Exploiting Subjectivity in Sketch-Based Person Retrieval
Kejun Lin, Zhixiang Wang, Zheng Wang +2
Person re-identification (re-ID) requires densely distributed cameras. In practice, the person of interest may not be captured by cameras and, therefore, needs to be retrieved usin…
cs.CV2024
Think-Program-reCtify: 3D Situated Reasoning with Large Language Models
Qingrong He, Kejun Lin, Shizhe Chen +2
This work addresses the 3D situated reasoning task which aims to answer questions given egocentric observations in a 3D environment. The task remains challenging as it requires com…