7 papers · 1 filter
GeoWorldAD: Geometry World Action Model for Autonomous Driving
Songyan Zhang, Jinyuan Tian, Hanbing Li +9
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual ob…
UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling
Tao Xu, Runhao Zhang, Zhijian Huang +7
Occluded tasks remain a bottleneck in robot manipulation. Existing solutions either deploy additional physical cameras requiring training-inference camera parity, or rely on explic…
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
Junli Wang, Zhihua Hua, Xueyi Liu +7
Existing imitation learning methods for end-to-end autonomous driving predominantly learn from successful demonstrations by minimizing geometric deviations from expert trajectories…
Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance
Lingfeng Zhang, Xiaoshuai Hao, Xizhou Bu +11
Assisting humans in open-world outdoor environments requires robots to translate high-level natural-language intentions into safe, long-horizon, and socially compliant navigation b…
Think before Go: Hierarchical Reasoning for Image-goal Navigation
Pengna Li, Kangyi Wu, Shaoqing Xu +5
Image-goal navigation steers an agent to a target location specified by an image in unseen environments. Existing methods primarily handle this task by learning an end-to-end navig…
Learning from Mistakes: Post-Training for Driving VLA with Takeover Data
Yinfeng Gao, Deqing Liu, Qichao Zhang +7
Current Vision-Language-Action (VLA) paradigms in end-to-end autonomous driving rely on offline training from static datasets, leaving them vulnerable to distribution shift. Recent…