2 papers
cs.CV2026
Perception-Aware Multimodal Spatial Reasoning from Monocular Images
Yanchun Cheng, Rundong Wang, Xulei Yang +4
Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, partic…
cs.CV2025
DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models
Xunjie He, Christina Dao Wen Lee, Meiling Wang +4
Collaborative perception plays a crucial role in enhancing environmental understanding by expanding the perceptual range and improving robustness against sensor failures, which pri…