activity
20242026
collaborators

8 papers

cs.RO2026

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

Haozhuang Chi, Daosheng Qiu, Hao Su +4

Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models forecast the external environment…

cs.RO2026

Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling

Daosheng Qiu, Haozhuang Chi, Hao Su +4

Continuous driver monitoring in automated vehicles requires low-latency inference while avoiding unsafe decisions under uncertain driver states. Large vision-language models provid…

cs.RO2026

SAGE-Nav: Leveraging LLM Planning and Alignment Fusion for Hierarchical Scene Graph-Guided Navigation

Hao Su, Yuehao Huang, Yukai Ma +2

Object-Goal Navigation (ObjNav) requires embodied agents to autonomously locate specified targets using only egocentric visual observations. Existing monolithic methods struggle wi…

cs.CV2026

DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

Jingke Wang, Zhenru Zhao, Shuangming Lei +8

Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow language guidances. However,…

cs.RO2025

Responsive Noise-Relaying Diffusion Policy: Responsive and Efficient Visuomotor Control

Zhuoqun Chen, Xiu Yuan, Tongzhou Mu +1

Imitation learning is an efficient method for teaching robots a variety of tasks. Diffusion Policy, which uses a conditional denoising diffusion process to generate actions, has de…

cs.LG2025

Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning

Adrià López Escoriza, Nicklas Hansen, Stone Tao +2

Long-horizon tasks in robotic manipulation present significant challenges in reinforcement learning (RL) due to the difficulty of designing dense reward functions and effectively e…