3 papers
cs.CV2026
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
Rui Zhao, Haofeng Hu, Zhenhai Gao +2
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalizatio…
cs.RO2026
ObjView-Bench: Rethinking Difficulty and Deployment for Object-Centric View Planning
Sicong Pan, Hao Hu, Xuying Huang +2
Object-centric view planning is a core component of active geometric 3D reconstruction in robotics, yet existing evaluations often conflate object complexity, planning difficulty,…
cs.CV2025
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning
Rui Zhao, Qirui Yuan, Jinyu Li +4
End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal La…