collaborators

6 papers

cs.LG2026

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning

Xuyang Chen, Keyu Yan, Guojian Wang +1

Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or cos…

cs.CL2026

OpenClaw-RL: Train Any Agent Simply by Talking

Yinjie Wang, Xuyang Chen, Xiaolong Jin +2

Every agent interaction generates a next-state signal, namely the user reply, tool output, terminal or GUI state change that follows each action, yet no existing agentic RL system…

cs.LG2025

Target-Aligned Fusion for Decision-Sequence Learning under Dynamics Shift

Guojian Wang, Quinson Hon, Xuyang Chen +1

External trajectories can improve offline decision-sequence learning, but dynamics shift may make some source subsequences inconsistent with the target environment. We study how to…

cs.LG2025

Preference-Guided Reinforcement Learning for Efficient Exploration

Guojian Wang, Jianxiang Liu, Xinyuan Li +4

In this paper, we investigate preference-based reinforcement learning (PbRL), which enables reinforcement learning (RL) agents to learn from human feedback. This is particularly va…

cs.LG2025

Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach

Xuyang Chen, Keyu Yan, Wenhan Cao +1

Offline reinforcement learning (RL) learns policies from fixed datasets without online interactions, but suffers from distribution shift, causing inaccurate evaluation and overesti…

cs.LG2025

Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

Xuyang Chen, Jingliang Duan, Lin Zhao

Actor-critic methods have achieved state-of-the-art performance in various challenging tasks. However, theoretical understandings of their performance remain elusive and challengin…