6 papers · 1 filter
Data Pyramid for Embodied Manipulation: A Survey
Yifan Ye, Yankai Fu, Yaoxu Lv +26
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…
DynamicWAM: Dual-Path Motion Conditioning for World-Action Models in Dynamic Manipulation
Yunfan Lou, Hewen Gao, Xiyu Zhu +6
Dynamic manipulation requires robots to infer target motion and respond promptly, yet existing World-Action Models (WAMs) typically condition only on the current frame and execute…
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation
Shuyi Zhang, Yunfan Lou, Hongyang Cheng +8
Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-tuning can surpass this limit…
Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination
Jiajun Li, Tiecheng Guo, Yifan Ye +9
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely o…
Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation
Yunfan Lou, Yifan Ye, Yankai Fu +7
World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on v…
Mask World Model: Predicting What Matters for Robust Robot Policy Learning
Yunfan Lou, Xiaowei Chi, Xiaojie Zhang +9
World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often fo…