45 citations · 138 across the 16 of their papers we have counts for
10 papers · 1 filter
Reinforcement learning with a network of spiking agents
Sneha Aenugu, Abhishek Sharma, Sasikiran Yelamarthi +3
Neuroscientific theory suggests that dopaminergic neurons broadcast global reward prediction errors to large areas of the brain influencing the synaptic plasticity of the neurons i…
Classical Policy Gradient: Preserving Bellman's Principle of Optimality
Philip S. Thomas, Scott M. Jordan, Yash Chandak +2
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the…
Is the Policy Gradient a Gradient?
Chris Nota, Philip S. Thomas
The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the di…
Reinforcement Learning When All Actions are Not Always Available
Yash Chandak, Georgios Theocharous, Blossom Metevier +1
The Markov decision process (MDP) formulation used to model many real-world sequential decision making problems does not efficiently capture the setting where the set of available…
Lifelong Learning with a Changing Action Set
Yash Chandak, Georgios Theocharous, Chris Nota +1
In many real-world sequential decision making problems, the number of available actions (decisions) can vary over time. While problems like catastrophic forgetting, changing transi…
A New Confidence Interval for the Mean of a Bounded Random Variable
Erik Learned-Miller, Philip S. Thomas
We present a new method for constructing a confidence interval for the mean of a bounded random variable from samples of the random variable. We conjecture that the confidence inte…