3 papers
cs.CV2026
Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering
Shaozu Ding, Linan Song, Dajiang Suo
Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic scenes. However, existing bench…
cs.CV2026
RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception
Shaozu Ding, Linan Song, Marco De Vincenzi +1
LiDAR has increasingly been integrated into traffic cameras to expand coverage and mitigate occlusion in roadside cooperative perception. However, how unimodal and camera-LiDAR fus…
cs.RO2024
High and Low Resolution Tradeoffs in Roadside Multimodal Sensing
Shaozu Ding, Yihong Tang, Marco De Vincenzi +1
Balancing cost and performance is crucial when choosing high- versus low-resolution point-cloud roadside sensors. For example, LiDAR delivers dense point cloud, while 4D millimeter…