45 citations · 144 across the 25 of their papers we have counts for
9 papers · 2 filters
Reinforcement learning with a network of spiking agents
Sneha Aenugu, Abhishek Sharma, Sasikiran Yelamarthi +3
Neuroscientific theory suggests that dopaminergic neurons broadcast global reward prediction errors to large areas of the brain influencing the synaptic plasticity of the neurons i…
Classical Policy Gradient: Preserving Bellman's Principle of Optimality
Philip S. Thomas, Scott M. Jordan, Yash Chandak +2
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the…
Is the Policy Gradient a Gradient?
Chris Nota, Philip S. Thomas
The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the di…
Reinforcement Learning When All Actions are Not Always Available
Yash Chandak, Georgios Theocharous, Blossom Metevier +1
The Markov decision process (MDP) formulation used to model many real-world sequential decision making problems does not efficiently capture the setting where the set of available…
Lifelong Learning with a Changing Action Set
Yash Chandak, Georgios Theocharous, Chris Nota +1
In many real-world sequential decision making problems, the number of available actions (decisions) can vary over time. While problems like catastrophic forgetting, changing transi…
A Meta-MDP Approach to Exploration for Lifelong Reinforcement Learning
Francisco M. Garcia, Philip S. Thomas
In this paper we consider the problem of how a reinforcement learning agent that is tasked with solving a sequence of reinforcement learning problems (a sequence of Markov decision…