9 papers
VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning
Tianxing Zhou, Feiyang Xue, Zhangchen Ye +3
Imitation learning is a prominent paradigm for robotic manipulation. However, existing visual imitation methods map 2D image observations directly to 3D action outputs, imposing a…
STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning
Zhihao Liu, Qiuyi Gu, Yitao Wang +16
Real-world robot learning increasingly relies on heterogeneous data, but demonstrations and rollouts often mix useful progress with stalls, corrections, and suboptimal behavior. Ef…
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
Zhennan Jiang, Shangqing Zhou, Yutong Jiang +11
Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requirement for massive real-world interact…
Human Universal Grasping
Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu +5
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from hum…
RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI
Hongzhi Zang, Shu'ang Yu, Hao Lin +14
Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitraril…
FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
Guangyan Chen, Meiling Wang, Te Cui +8
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in foundation models, particularly Vis…