5 papers
Learning Reusable Options for Multi-Task Reinforcement Learning
Francisco M. Garcia, Chris Nota, Philip S. Thomas
Reinforcement learning (RL) has become an increasingly active area of research in recent years. Although there are many algorithms that allow an agent to solve tasks efficiently, t…
Classical Policy Gradient: Preserving Bellman's Principle of Optimality
Philip S. Thomas, Scott M. Jordan, Yash Chandak +2
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the…
Is the Policy Gradient a Gradient?
Chris Nota, Philip S. Thomas
The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the di…
Lifelong Learning with a Changing Action Set
Yash Chandak, Georgios Theocharous, Chris Nota +1
In many real-world sequential decision making problems, the number of available actions (decisions) can vary over time. While problems like catastrophic forgetting, changing transi…
Asynchronous Coagent Networks
James E. Kostas, Chris Nota, Philip S. Thomas
Coagent policy gradient algorithms (CPGAs) are reinforcement learning algorithms for training a class of stochastic neural networks called coagent networks. In this work, we prove…