collaborators

5 papers

cs.RO2026

VINE: Taming Generative Control Policies for Reinforcement Learning

Rushuai Yang, Zhuo Han, Houlin Li +10

Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of…

cs.RO2026

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training

Rushuai Yang, Hecheng Wang, Zhichao Wu +11

We study how to improve large foundation vision-language-action (VLA) systems through human-in-the-loop reinforcement learning (RL) in real-world environments. A key challenge is l…

cs.AI2026

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets

Yi Chen, Rushuai Yang, Qiang Chen +2

Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operational constraints. These feature…

cs.LG2025

Unsupervised Skill Discovery through Skill Regions Differentiation

Ting Xiao, Jiakun Zheng, Rushuai Yang +4

Unsupervised Reinforcement Learning (RL) aims to discover diverse behaviors that can accelerate the learning of downstream tasks. Previous methods typically focus on entropy-based…

cs.CL2025

Supervised Optimism Correction: Be Confident When LLMs Are Sure

Junjie Zhang, Rushuai Yang, Shunyu Liu +5

In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing…