24 citations · 32 across the 9 of their papers we have counts for
1 paper · 1 filter
Phillip Swazinna, Steffen Udluft, Daniel Hein +1
In offline reinforcement learning, a policy needs to be learned from a single pre-collected dataset. Typically, policies are thus regularized during training to behave similarly to…