1 paper
Shin-ichi Maeda, Hayato Watahiki, Shintarou Okada +1
Practical reinforcement learning problems are often formulated as constrained Markov decision process (CMDP) problems, in which the agent has to maximize the expected return while…