8 papers
Learning 4D Geometric Priors for Inference-Efficient World Action Models
Jianjun Zhang, Jian Zhu, Taiyi Su +4
World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-…
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
Jian Zhu, Jianjun Zhang, Taiyi Su +10
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learni…
PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation
Chong Ma, Taiyi Su, Jian Zhu +4
Vision-language-action (VLA) policies operate in a closed loop in real-world robot tasks: a robot observes the scene, executes an action chunk, and conditions its next decision on…
HumanoidExo: Scalable Whole-Body Humanoid Manipulation via Wearable Exoskeleton
Rui Zhong, Yizhe Sun, Junjie Wen +7
A significant bottleneck in humanoid policy learning is the acquisition of large-scale, diverse datasets, as collecting reliable real-world data remains both difficult and cost-pro…
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
Qiyuan Zeng, Chengmeng Li, Jude St. John +5
We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples…
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
Junjie Wen, Minjie Zhu, Jiaming Liu +6
Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought t…