6 papers
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
Xuyang Chen, Keyu Yan, Guojian Wang +1
Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or cos…
OpenClaw-RL: Train Any Agent Simply by Talking
Yinjie Wang, Xuyang Chen, Xiaolong Jin +2
Every agent interaction generates a next-state signal, namely the user reply, tool output, terminal or GUI state change that follows each action, yet no existing agentic RL system…
Target-Aligned Fusion for Decision-Sequence Learning under Dynamics Shift
Guojian Wang, Quinson Hon, Xuyang Chen +1
External trajectories can improve offline decision-sequence learning, but dynamics shift may make some source subsequences inconsistent with the target environment. We study how to…
Preference-Guided Reinforcement Learning for Efficient Exploration
Guojian Wang, Jianxiang Liu, Xinyuan Li +4
In this paper, we investigate preference-based reinforcement learning (PbRL), which enables reinforcement learning (RL) agents to learn from human feedback. This is particularly va…
Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach
Xuyang Chen, Keyu Yan, Wenhan Cao +1
Offline reinforcement learning (RL) learns policies from fixed datasets without online interactions, but suffers from distribution shift, causing inaccurate evaluation and overesti…
Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator
Xuyang Chen, Jingliang Duan, Lin Zhao
Actor-critic methods have achieved state-of-the-art performance in various challenging tasks. However, theoretical understandings of their performance remain elusive and challengin…