collaborators

6 papers

cs.CV2026

UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

Yongkang Li, Lijun Zhou, Sixu Yan +11

Vision-Language-Action (VLA) models have recently emerged in autonomous driving, with the promise of leveraging rich world knowledge to improve the cognitive capabilities of drivin…

cs.RO2026

OmniTrack: General Motion Tracking via Physics-Consistent Reference

Yuhan Li, Peiyuan Zhi, Yunshen Wang +6

Learning motion tracking from rich human motion data is a foundational task for achieving general control in humanoid robots, enabling them to perform diverse behaviors. However, d…

cs.CV2025

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

Yongkang Li, Kaixin Xiong, Xiangyu Guo +12

Recent studies have explored leveraging the world knowledge and cognitive capabilities of Vision-Language Models (VLMs) to address the long-tail problem in end-to-end autonomous dr…

cs.RO2025

M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes

Zeyu Zhang, Sixu Yan, Muzhi Han +4

We propose M3Bench, a new benchmark for whole-body motion generation in mobile manipulation tasks. Given a 3D scene context, M3Bench requires an embodied agent to reason about its…

cs.CV2025

DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

Bencheng Liao, Shaoyu Chen, Haoran Yin +8

Recently, the diffusion model has emerged as a powerful generative technique for robotic policy learning, capable of modeling multi-mode action distributions. Leveraging its capabi…

cs.CV2025

Towards Fast, Memory-based and Data-Efficient Vision-Language Policy

Haoxuan Li, Sixu Yan, Yuhan Li +1

Vision Language Models (VLMs) pretrained on Internet-scale vision-language data have demonstrated the potential to transfer their knowledge to robotic learning. However, the existi…