1 paper
Zifan Wu, Bo Tang, Qian Lin +5
Primal-dual safe RL methods commonly perform iterations between the primal update of the policy and the dual update of the Lagrange Multiplier. Such a training paradigm is highly s…