3 papers
cs.CL2026
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
Encheng Su, Jianyu Wu, Jinouwen Zhang +9
Long-horizon memory question answering often requires sparse evidence from heterogeneous histories, including events, object states, visual observations, temporal relations, and ca…
cs.CV2026
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Qiucheng Yu, Ruijie Xu, Mingang Chen +2
Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing benchmarks suffer from three fundamental…
cs.CV2025
SHTOcc: Effective 3D Occupancy Prediction with Sparse Head and Tail Voxels
Qiucheng Yu, Yuan Xie, Xin Tan
3D occupancy prediction has attracted much attention in the field of autonomous driving due to its powerful geometric perception and object recognition capabilities. However, exist…