13 papers
SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model
Zewei Zhou, Ruining Yang, Xuewei +8
Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. Howe…
BridgeSim: Unveiling the OL-CL Gap in End-to-End Autonomous Driving
Seth Z. Zhao, Luobin Wang, Hongwei Ruan +13
Open-loop (OL) to closed-loop (CL) gap (OL-CL gap) exists when OL-pretrained policies scoring high in OL evaluations fail to transfer effectively in closed-loop (CL) deployment. In…
MDG: Masked Denoising Generation for Multi-Agent Behavior Modeling in Traffic Environments
Zhiyu Huang, Zewei Zhou, Tianhui Cai +2
Modeling realistic and interactive multi-agent behavior is critical to autonomous driving and traffic simulation. However, existing diffusion and autoregressive approaches are limi…
MIC-BEV: Multi-Infrastructure Camera Bird's-Eye-View Transformer with Relation-Aware Fusion for 3D Object Detection
Yun Zhang, Zhaoliang Zheng, Johnson Liu +5
Infrastructure-based perception plays a crucial role in intelligent transportation systems, offering global situational awareness and enabling cooperative autonomy. However, existi…
RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes
Xinyi Liu, Mohammadreza Fani Sani, Zewei Zhou +3
Despite rapid progress in autonomous robotics, executing complex or long-horizon tasks remains a fundamental challenge. Most current approaches follow an open-loop paradigm with li…
TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and Prediction
Zewei Zhou, Seth Z. Zhao, Tianhui Cai +3
End-to-end training of multi-agent systems offers significant advantages in improving multi-task performance. However, training such models remains challenging and requires extensi…