collaborators

22 papers

cs.RO2026

DreamWAM: Beyond RGB Future Prediction for World Action Models

Shanglin Yuan, Weiheng Zhao, Xin Shi +6

World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…

cs.RO2026

Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

Xufeng Zhao, Fuzhi Yang, Jianhui Chen +17

The motion controller is one of the most fundamental modules in embodied intelligence systems. Driven by large-scale human motion-capture data and the motion-tracking paradigm, hum…

cs.RO2026

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

Shanglin Yuan, Weiheng Zhao, Xianda Guo +4

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiot…

cs.CV2026

VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

Bo Jiang, Shaoyu Chen, Hao Gao +4

Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of planning make it challenging. Existin…

cs.CV2026

RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

Hao Gao, Shaoyu Chen, Yifan Zhu +4

High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interactions. Although diffusion-ba…

cs.CV2026

UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

Yongkang Li, Lijun Zhou, Sixu Yan +11

Vision-Language-Action (VLA) models have recently emerged in autonomous driving, with the promise of leveraging rich world knowledge to improve the cognitive capabilities of drivin…