collaborators

8 papers

cs.RO2026

Learning 4D Geometric Priors for Inference-Efficient World Action Models

Jianjun Zhang, Jian Zhu, Taiyi Su +4

World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-…

cs.RO2026

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

Jian Zhu, Jianjun Zhang, Taiyi Su +10

World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learni…

cs.RO2026

PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation

Chong Ma, Taiyi Su, Jian Zhu +4

Vision-language-action (VLA) policies operate in a closed loop in real-world robot tasks: a robot observes the scene, executes an action chunk, and conditions its next decision on…

cs.RO2025

HumanoidExo: Scalable Whole-Body Humanoid Manipulation via Wearable Exoskeleton

Rui Zhong, Yizhe Sun, Junjie Wen +7

A significant bottleneck in humanoid policy learning is the acquisition of large-scale, diverse datasets, as collecting reliable real-world data remains both difficult and cost-pro…

cs.RO2025

ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations

Qiyuan Zeng, Chengmeng Li, Jude St. John +5

We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples…

cs.RO2025

dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought

Junjie Wen, Minjie Zhu, Jiaming Liu +6

Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought t…