22 citations · 22 across the 2 of their papers we have counts for
3 papers
Classical Policy Gradient: Preserving Bellman's Principle of Optimality
Philip S. Thomas, Scott M. Jordan, Yash Chandak +2
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the…
Asynchronous Coagent Networks
James E. Kostas, Chris Nota, Philip S. Thomas
Coagent policy gradient algorithms (CPGAs) are reinforcement learning algorithms for training a class of stochastic neural networks called coagent networks. In this work, we prove…
Learning Action Representations for Reinforcement Learning
Yash Chandak, Georgios Theocharous, James Kostas +2
Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the str…