8 papers
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Chaohong Guo, Xun Mo, Yongwei Nie +3
Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforc…
Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models
Guo Li, Weihong Chen, Yongfu Fan
Diffusion models have demonstrated powerful performance in generating high-quality images. A typical example is text-to-image generator like Stable Diffusion. However, their widesp…
Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability
Zhaoyu Chen, Hongnan Lin, Yongwei Nie +4
Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods…
Beyond Inference Intervention: Identity-Decoupled Diffusion for Face Anonymization
Haoxin Yang, Yihong Lin, Jingdan Kang +4
Face anonymization aims to conceal identity information while preserving non-identity attributes. Mainstream diffusion models rely on inference-time interventions such as negative…
Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection
Yuyang Yu, Zhengwei Chen, Xuemiao Xu +4
3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based…
RecDreamer: Consistent Text-to-3D Generation via Uniform Score Distillation
Chenxi Zheng, Yihong Lin, Bangzhen Liu +3
Current text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. Thi…