activity
20232026
collaborators

5 papers

cs.CV2026

VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving

Rui Zhao, Jianlin Yu, Zhenhai Gao +2

End-to-end autonomous driving requires models to understand traffic scenes, infer driving intent, and generate executable motion plans. Recent vision-language-action (VLA) models i…

cs.CV2026

VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving

Rui Zhao, Haofeng Hu, Zhenhai Gao +2

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalizatio…

cs.RO2025

DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy

Rui Zhao, Yuze Fan, Ziguo Chen +2

End-to-end learning has emerged as a transformative paradigm in autonomous driving. However, the inherently multimodal nature of driving behaviors and the generalization challenges…

cs.CV2025

Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning

Rui Zhao, Qirui Yuan, Jinyu Li +4

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal La…

cs.CV2023

An Enhanced Low-Resolution Image Recognition Method for Traffic Environments

Zongcai Tan, Zhenhai Gao

Currently, low-resolution image recognition is confronted with a significant challenge in the field of intelligent traffic perception. Compared to high-resolution images, low-resol…