activity
20162021
most citedPolicy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

45 citations · 138 across the 16 of their papers we have counts for

collaborators
Showing 2019Show all

10 papers · 1 filter

cs.LG20193 cited

Reinforcement learning with a network of spiking agents

Sneha Aenugu, Abhishek Sharma, Sasikiran Yelamarthi +3

Neuroscientific theory suggests that dopaminergic neurons broadcast global reward prediction errors to large areas of the brain influencing the synaptic plasticity of the neurons i…

cs.LG2019

Classical Policy Gradient: Preserving Bellman's Principle of Optimality

Philip S. Thomas, Scott M. Jordan, Yash Chandak +2

We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the…

cs.LG2019

Is the Policy Gradient a Gradient?

Chris Nota, Philip S. Thomas

The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the di…

cs.LG2019

Reinforcement Learning When All Actions are Not Always Available

Yash Chandak, Georgios Theocharous, Blossom Metevier +1

The Markov decision process (MDP) formulation used to model many real-world sequential decision making problems does not efficiently capture the setting where the set of available…

cs.LG2019

Lifelong Learning with a Changing Action Set

Yash Chandak, Georgios Theocharous, Chris Nota +1

In many real-world sequential decision making problems, the number of available actions (decisions) can vary over time. While problems like catastrophic forgetting, changing transi…

math.ST2019

A New Confidence Interval for the Mean of a Bounded Random Variable

Erik Learned-Miller, Philip S. Thomas

We present a new method for constructing a confidence interval for the mean of a bounded random variable from samples of the random variable. We conjecture that the confidence inte…