2 papers
cs.LG2025
Unsupervised Skill Discovery through Skill Regions Differentiation
Ting Xiao, Jiakun Zheng, Rushuai Yang +4
Unsupervised Reinforcement Learning (RL) aims to discover diverse behaviors that can accelerate the learning of downstream tasks. Previous methods typically focus on entropy-based…
cs.CL2025
Supervised Optimism Correction: Be Confident When LLMs Are Sure
Junjie Zhang, Rushuai Yang, Shunyu Liu +5
In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing…