1 paper
Qiuhao Wang, Shaohang Xu, Chin Pang Ho +1
We develop a generic policy gradient method with the global optimality guarantee for robust Markov Decision Processes (MDPs). While policy gradient methods are widely used for solv…