4 papers
CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning
Houlin Li, Minghui Xu, Guo Xu +8
Human-in-the-loop real-world reinforcement learning enables rapid acquisition of effective robotic manipulation policies for individual tasks, often within tens of minutes. Yet it…
VINE: Taming Generative Control Policies for Reinforcement Learning
Rushuai Yang, Zhuo Han, Houlin Li +10
Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of…
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
Rushuai Yang, Hecheng Wang, Zhichao Wu +11
We study how to improve large foundation vision-language-action (VLA) systems through human-in-the-loop reinforcement learning (RL) in real-world environments. A key challenge is l…
Continually Evolving Skill Knowledge in Vision Language Action Model
Yuxuan Wu, Guangming Wang, Zhiheng Yang +4
Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains challenging, especially for efficient adap…