7 papers
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
Yiru Wang, Zichong Gu, Yu Gao +5
Vision-Language-Action (VLA) models offer promising capabilities for autonomous driving through multimodal understanding. However, their utilization in safety-critical scenarios is…
DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
Yu Gao, Anqing Jiang, Yiru Wang +7
Conventional end-to-end (E2E) driving models are effective at generating physically plausible trajectories, but often fail to generalize to long-tail scenarios due to the lack of e…
FlowDrive: Energy Flow Field for End-to-End Autonomous Driving
Hao Jiang, Zhipeng Zhang, Yu Gao +11
Recent advances in end-to-end autonomous driving leverage multi-view images to construct BEV representations for motion planning. In motion planning, autonomous vehicles need consi…
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
Anqing Jiang, Yu Gao, Yiru Wang +11
Vision-Language-Action (VLA) models have demonstrated potential in autonomous driving. However, two critical challenges hinder their development: (1) Existing VLA architectures are…
DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online HD Map Diffusion
Zhigang Sun, Yiru Wang, Anqing Jiang +13
Autonomous driving requires accurate scene understanding, including road geometry, traffic agents, and their semantic relationships. In online HD map generation scenarios, raster-b…
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
Anqing Jiang, Yu Gao, Zhigang Sun +11
Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which ena…