4 papers · 1 filter
General Flexible -divergence for Challenging Offline RL Datasets with Low Stochasticity and Diverse Behavior Policies
Jianxun Wang, Grant C. Forbes, Leonardo Villalobos-Arias +1
Offline RL algorithms aim to improve upon the behavior policy that produces the collected data while constraining the learned policy to be within the support of the dataset. Howeve…
Action-Dependent Optimality-Preserving Reward Shaping
Grant C. Forbes, Jianxun Wang, Leonardo Villalobos-Arias +2
Recent RL research has utilized reward shaping--particularly complex shaping rewards such as intrinsic motivation (IM)--to encourage agent exploration in sparse-reward environments…
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
Grant C. Forbes, Leonardo Villalobos-Arias, Jianxun Wang +2
Recently there has been a proliferation of intrinsic motivation (IM) reward-shaping methods to learn in complex and sparse-reward environments. These methods can often inadvertentl…
Potential-Based Reward Shaping For Intrinsic Motivation
Grant C. Forbes, Nitish Gupta, Leonardo Villalobos-Arias +3
Recently there has been a proliferation of intrinsic motivation (IM) reward-shaping methods to learn in complex and sparse-reward environments. These methods can often inadvertentl…