2 papers
cs.LG2020
Logistic Q-Learning
Joan Bas-Serrano, Sebastian Curi, Andreas Krause +1
We propose a new reinforcement learning algorithm derived from a regularized linear-programming formulation of optimal control in MDPs. The method is closely related to the classic…
math.OC2019
Faster saddle-point optimization for solving large-scale Markov decision processes
Joan Bas-Serrano, Gergely Neu
We consider the problem of computing optimal policies in average-reward Markov decision processes. This classical problem can be formulated as a linear program directly amenable to…