Publications (5)
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
Encheng Su, Jianyu Wu, Jinouwen Zhang +9
Long-horizon memory question answering often requires sparse evidence from heterogeneous histories, including events, object states, visual observations, temporal relations, and ca…
SHTOcc: Effective 3D Occupancy Prediction with Sparse Head and Tail Voxels
Qiucheng Yu, Yuan Xie, Xin Tan
3D occupancy prediction has attracted much attention in the field of autonomous driving due to its powerful geometric perception and object recognition capabilities. However, exist…
Human-Imperceptible Physical Adversarial Attack for NIR Face Recognition Models
Songyan Xie, Jinghang Wen, Encheng Su +1
Near-infrared (NIR) face recognition systems, which can operate effectively in low-light conditions or in the presence of makeup, exhibit vulnerabilities when subjected to physical…
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Qiucheng Yu, Ruijie Xu, Mingang Chen +2
The paper introduces TSHA, a large benchmark of real-world indoor safety hazard assessment questions for evaluating vision‑language models, and shows that training on this data imp…
Hiding in Plain Sight: An Effective Physical Adversarial Patch Attack against Visual-Infrared Fused Face Detection
Qiucheng Yu, Tao Ni, Yihe Zhou +2
Deep learning-based visual-infrared fused face detection models are increasingly deployed across a wide range of applications, yet they remain susceptible to adversarial patch atta…