2 papers
cs.CV2026
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Chaohong Guo, Xun Mo, Yongwei Nie +3
Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforc…
cs.CV2026
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
Chaohong Guo, Yihan He, Yongwei Nie +3
Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dyn…