activity
20242026
collaborators

10 papers

cs.CV2026

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

Chaohong Guo, Xun Mo, Yongwei Nie +3

Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforc…

cs.CV2026

T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding

Chaohong Guo, Yihan He, Yongwei Nie +3

Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dyn…

cs.AI2026

Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability

Zhaoyu Chen, Hongnan Lin, Yongwei Nie +4

Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods…

cs.CV2025

Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection

Yuyang Yu, Zhengwei Chen, Xuemiao Xu +4

3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based…

cs.CV2025

Enhancing Long Video Question Answering with Scene-Localized Frame Grouping

Xuyi Yang, Wenhao Zhang, Hongbo Jin +5

Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video…

cs.CV2025

RecDreamer: Consistent Text-to-3D Generation via Uniform Score Distillation

Chenxi Zheng, Yihong Lin, Bangzhen Liu +3

Current text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. Thi…