2 papers
cs.CV2026
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
Xiaowei Gao, Pengxiang Li, Yitai Cheng +4
Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off betw…
cs.CV2026
A Contrastive Learning Framework Empowered by Attention-based Feature Adaptation for Street-View Image Classification
Qi You, Yitai Cheng, Zichao Zeng +1
Street-view image attribute classification is a vital downstream task of image classification, enabling applications such as autonomous driving, urban analytics, and high-definitio…