1 paper · 1 filter
Ilias Kazantzidis, Timothy J. Norman, Yali Du +1
We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward functi…