1 paper
Hengrui Zhang, Youfang Lin, Sheng Han +2
Safety exploration can be regarded as a constrained Markov decision problem where the expected long-term cost is constrained. Previous off-policy algorithms convert the constrained…