1 paper · 1 filter
Vincent Corlay, Jean-Christophe Sibel
Standard Markov decision process (MDP) and reinforcement learning algorithms optimize the policy with respect to the expected gain. We propose an algorithm which enables to optimiz…