1 citations · 2 across the 3 of their papers we have counts for
3 papers
TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning
Batıkan Bora Ormancı, Phillip Swazinna, Steffen Udluft +1
In this paper, we investigate offline reinforcement learning (RL) with the goal of training a single robust policy that generalizes effectively across environments with unseen dyna…
Why long model-based rollouts are no reason for bad Q-value estimates
Philipp Wissmann, Daniel Hein, Steffen Udluft +1
This paper explores the use of model-based offline reinforcement learning with long model rollouts. While some literature criticizes this approach due to compounding errors, many p…
Safe Policy Improvement Approaches and their Limitations
Philipp Scholl, Felix Dietrich, Clemens Otte +1
Safe Policy Improvement (SPI) is an important technique for offline reinforcement learning in safety critical applications as it improves the behavior policy with a high probabilit…