1 paper
Xiaocheng Li, Huaiyang Zhong, Margaret L. Brandeau
The goal of a traditional Markov decision process (MDP) is to maximize expected cumulative reward over a defined horizon (possibly infinite). In many applications, however, a decis…