collaborators

11 papers

cs.RO2026

LaST: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

Zhuoyang Liu, Jiaming Liu, Hao Chen +11

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future obs…

cs.RO2026

URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets

Zhuangzhe Wu, Yue Xin, Chengkai Hou +4

Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, recovering them from visual observations is inherently chall…

cs.RO2026

Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation

Jingyang He, Guangrun Li, Jieyu Zhang +3

Robotic imitation learning is often treated as reproducing demonstrated actions, but actions are inherently embodiment-specific. When demonstrations come from humans or robots with…

cs.RO2026

HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Shuanghao Bai, Meng Li, Xinyuan Lv +14

Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making hi…

cs.RO2026

RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence

Chengkai Hou, Kun Wu, Jiaming Liu +30

While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstration…

cs.RO2026

SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System

Hao Wang, Chengkai Hou, Xianglong Li +7

Learning to control high-speed objects in dynamic environments represents a fundamental challenge in robotics. Table tennis serves as an ideal testbed for advancing robotic capabil…