11 papers
Data Pyramid for Embodied Manipulation: A Survey
Yifan Ye, Yankai Fu, Yaoxu Lv +26
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…
Hierarchical Denoising For Multi-Step Visual Reasoning
Zezhong Qian, Xiaowei Chi, Chak-Wing Mak +9
The paper introduces HDR, a hierarchical denoising framework for causal video generation that enables multi-step visual reasoning with low-latency streaming, achieving higher succe…
WAM4D: Fast 4D World Action Model via Spatial Register Tokens
Ying Li, Xiaobao Wei, Jiajun Cao +10
World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video o…
Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation
Yunfan Lou, Yifan Ye, Yankai Fu +7
World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on v…
TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation
Qinwen Xu, Jiaming Liu, Rui Zhou +11
Despite strong generalization capabilities, Vision-Language-Action (VLA) models remain constrained by the high cost of expert demonstrations and limited real-world interaction. Whi…
ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution
Chuyao Fu, Shengzhe Gan, Zhuoli Ouyang +5
End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, and such reactive planning ca…