collaborators

11 papers

cs.RO2026

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Shichao Fan, Kun Wu, Zhengping Che +12

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still f…

cs.RO2026

IOI: Decoupling Kinematics and Physics for Interactive World Models

Chengyu Bai, Peidong Jia, Tiecheng Guo +11

Developing generalist embodied agents requires interactive environments providing visually realistic feedback and accurate action-conditioned dynamics. Interactive world models add…

cs.RO2026

MV-WAM: Manifold-Aware World Action Model with Value Augmentation

Jintao Chen, Peidong Jia, Qingpo Wuwu +13

Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domai…

cs.RO2026

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

Kai Tang, Peidong Jia, Zhong Chu +15

Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However, existing safe reinforcement…

cs.RO2026

LaST: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

Zhuoyang Liu, Jiaming Liu, Hao Chen +11

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future obs…

cs.RO2026

Heracles: Bridging Precise Tracking and Generative Synthesis for General Humanoid Control

Zelin Tao, Zeran Su, Peiran Liu +13

Achieving general-purpose humanoid control requires a delicate balance between the precise execution of commanded motions and the flexible, anthropomorphic adaptability needed to r…