1 paper
Wensong Bai, Chao Zhang, Qihang Xu +3
Offline safe reinforcement learning (RL) aims to learn policies from a fixed dataset while maximizing performance under cumulative cost constraints. In practice, deployment require…