6 papers
DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions
Weicheng Zheng, Yixin Huang, Qiao Sun +2
Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise. We propose DriveMA, a Driv…
DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions
Weicheng Zheng, Yixin Huang, Qiao Sun +2
Driving Vision-Language-Action Models (Driving VLAs) commonly introduce natural-language reasoning as an intermediate interface for end-to-end planning, but reasoning-centric inter…
Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback
Derun Li, Changye Li, Yue Wang +9
Generating human-like and adaptive trajectories is essential for autonomous driving in dynamic environments. While generative models have shown promise in synthesizing feasible tra…
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
Ruixun Liu, Lingyu Kong, Derun Li +1
Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driv…
VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion
Shaoting Zhu, Linzhan Mou, Derun Li +3
Recent success in legged robot locomotion is attributed to the integration of reinforcement learning and physical simulators. However, these policies often encounter challenges whe…
SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
Shaoting Zhu, Derun Li, Linzhan Mou +3
The application of vision-language models (VLMs) has achieved impressive success in various robotics tasks. However, there are few explorations for these foundation models used in…