6 papers
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
Simon Sinong Zhan, Qingyuan Wu, Philip Wang +4
Offline-to-online deployment of reinforcement-learning (RL) agents must bridge two gaps: (1) the sim-to-real gap, where real systems add latency and other imperfections not present…
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
Simon Sinong Zhan, Philip Wang, Qingyuan Wu +4
In this paper, we aim to tackle the limitation of the Adversarial Inverse Reinforcement Learning (AIRL) method in stochastic environments where theoretical results cannot hold and…
A Unified Framework for Rethinking Policy Divergence Measures in GRPO
Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan +6
Reinforcement Learning with Verified Reward (RLVR) has emerged as a critical paradigm for advancing the reasoning capabilities of Large Language Models (LLMs). Most existing RLVR m…
Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning
Yuhui Wang, Qingyuan Wu, Dylan R. Ashley +4
The Value Iteration Network (VIN) is an end-to-end differentiable neural network architecture for planning. It exhibits strong generalization to unseen domains by incorporating a d…
Directly Forecasting Belief for Reinforcement Learning with Delays
Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan +6
Reinforcement learning (RL) with delays is challenging as sensory perceptions lag behind the actual events: the RL agent needs to estimate the real state of its environment based o…
Inverse Delayed Reinforcement Learning
Simon Sinong Zhan, Qingyuan Wu, Zhian Ruan +6
Inverse Reinforcement Learning (IRL) has demonstrated effectiveness in a variety of imitation tasks. In this paper, we introduce an IRL framework designed to extract rewarding feat…