4 papers
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Chaohong Guo, Xun Mo, Yongwei Nie +3
Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforc…
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
Chaohong Guo, Yihan He, Yongwei Nie +3
Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dyn…
Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability
Zhaoyu Chen, Hongnan Lin, Yongwei Nie +4
Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods…
DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering
Rongjia Zheng, Qing Zhang, Chengjiang Long +1
Recent methods have shown that pre-trained diffusion models can be fine-tuned to enable generative inverse rendering by learning image-conditioned noise-to-intrinsic mapping. Despi…