5 papers
VINE: Taming Generative Control Policies for Reinforcement Learning
Rushuai Yang, Zhuo Han, Houlin Li +10
Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of…
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
Rushuai Yang, Hecheng Wang, Zhichao Wu +11
We study how to improve large foundation vision-language-action (VLA) systems through human-in-the-loop reinforcement learning (RL) in real-world environments. A key challenge is l…
Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets
Yi Chen, Rushuai Yang, Qiang Chen +2
Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operational constraints. These feature…
Unsupervised Skill Discovery through Skill Regions Differentiation
Ting Xiao, Jiakun Zheng, Rushuai Yang +4
Unsupervised Reinforcement Learning (RL) aims to discover diverse behaviors that can accelerate the learning of downstream tasks. Previous methods typically focus on entropy-based…
Supervised Optimism Correction: Be Confident When LLMs Are Sure
Junjie Zhang, Rushuai Yang, Shunyu Liu +5
In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing…