6 papers
Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation
Tianshuai Hu, Yangyi Zhong, Zeying Gong +7
Vision-Language Navigation in dynamic, human-centric environments exposes a fundamental tension: linguistic reasoning is slow and deliberative, whereas safe, socially compliant pla…
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
Xiaodong Mei, Diankun Zhang, Hongwei Xie +3
Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, whi…
NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation
Tianshuai Hu, Zeying Gong, Lingdong Kong +7
Social navigation requires robots to act safely in dynamic human environments. Effective behavior demands thinking ahead: reasoning about how the scene and pedestrians evolve under…
NetRoller: Interfacing General and Specialized Models for End-to-End Autonomous Driving
Ren Xin, Hongji Liu, Xiaodong Mei +4
Integrating General Models (GMs) such as Large Language Models (LLMs), with Specialized Models (SMs) in autonomous driving tasks presents a promising approach to mitigating challen…
HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning
Xiaodong Mei, Sheng Wang, Jie Cheng +2
Motion forecasting represents a critical challenge in autonomous driving systems, requiring accurate prediction of surrounding agents' future trajectories. While existing approache…
LHPF: Look back the History and Plan for the Future in Autonomous Driving
Sheng Wang, Yao Tian, Xiaodong Mei +5
Decision-making and planning in autonomous driving critically reflect the safety of the system, making effective planning imperative. Current imitation learning-based planning algo…