1 paper
Xingtu Liu, Lin F. Yang, Sharan Vaswani
We consider infinite-horizon γ-discounted (linear) constrained Markov decision processes (CMDPs) where the objective is to find a policy that maximizes the expected cumulative re…