12 papers
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation
Shuyi Zhang, Yunfan Lou, Hongyang Cheng +8
Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-tuning can surpass this limit…
Factor-Aware Mixture-of-Experts with Pretrained Encoder for Combinatorial Generalization
Feihong Zhang, Guojian Zhan, Zeyu He +8
The integration of pretrained encoders with diffusion policies has become a dominant paradigm for visual robotic manipulation. However, it still struggles to generalize across comp…
Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation
Yunfan Lou, Yifan Ye, Yankai Fu +7
World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on v…
URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets
Zhuangzhe Wu, Yue Xin, Chengkai Hou +4
Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, recovering them from visual observations is inherently chall…
Mask World Model: Predicting What Matters for Robust Robot Policy Learning
Yunfan Lou, Xiaowei Chi, Xiaojie Zhang +9
World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often fo…
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
Guangrun Li, Yaoxu Lyu, Zhuoyang Liu +3
Large-scale pre-training using egocentric human videos has proven effective for robot learning. However, the models pre-trained on such data can be suboptimal for robot learning du…