collaborators

12 papers

cs.RO2026

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies

Jiarui Yang, Yehao Lu, Yuning Su +9

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that joint…

cs.RO2026

Perceptive Behavior Foundation Model: Adapting Human Motion Priors to Robot-Centric Terrain

Zifan Wang, Yizhao Li, Teli Ma +5

Humanoid behavior foundation models aim to acquire reusable whole-body control policies from broad human motion priors, enabling a single controller to produce diverse and expressi…

cs.RO2026

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

Jia Zheng, Teli Ma, Yudong Fan +3

World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional vi…

cs.RO2026

DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control

Teli Ma, Jia Zheng, Zifan Wang +4

Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretrainin…

cs.CV2026

EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations

Jiayi Liu, Jiaming Zhou, Ke Ye +3

Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless…

cs.RO2025

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

Ke Ye, Jiaming Zhou, Yuanfeng Qiu +4

Generalizing to long-horizon manipulation tasks in a zero-shot setting remains a central challenge in robotics. Current multimodal foundation based approaches, despite their capabi…