From the 1 of 15 linked papers with an AI index.
15 papers
Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching
Runwei Guan, Di Tian, Ningwei Ouyang +9
As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor e…
4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception
Xiaokai Bai, Lianqing Zheng, Runwei Guan +3
The paper introduces a 4D radar‑camera framework that jointly performs 3D object detection and dense occupancy prediction by modeling occupancy as a persistent scene state and usin…
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos
Leyuan Yu, Xiao Tang, Minghao Liu +6
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically v…
RC-GeoCP: Geometric Consensus for Radar-Camera Collaborative Perception
Xiaokai Bai, Lianqing Zheng, Runwei Guan +3
Collaborative perception (CP) improves scene understanding through multi-agent information sharing, yet LiDAR-centric systems remain costly and vulnerable in adverse weather. Camer…
StreamPPG: Low-Latency rPPG Estimation via Consistent Privileged Learning
Yiming Li, Yihan Yang, Yuguang Chu +6
Remote photoplethysmography (rPPG) estimates the blood volume pulse (BVP) signal from facial videos, enabling contact-free health monitoring. Conventional clip-wise approaches, whi…
SD4R: Sparse-to-Dense Learning for 3D Object Detection with 4D Radar
Xiaokai Bai, Jiahao Cheng, Songkai Wang +5
4D radar measurements offer an affordable and weather-robust solution for 3D perception. However, the inherent sparsity and noise of radar point clouds present significant challeng…