activity
20162026
most citedPolicy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

45 citations · 144 across the 25 of their papers we have counts for

collaborators
Showing 2019 · cs.LGShow all

9 papers · 2 filters

cs.LG2019★ 3 cited

Reinforcement learning with a network of spiking agents

Sneha Aenugu, Abhishek Sharma, Sasikiran Yelamarthi +3

Neuroscientific theory suggests that dopaminergic neurons broadcast global reward prediction errors to large areas of the brain influencing the synaptic plasticity of the neurons i…

cs.LG2019

Classical Policy Gradient: Preserving Bellman's Principle of Optimality

Philip S. Thomas, Scott M. Jordan, Yash Chandak +2

We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the…

cs.LG2019

Is the Policy Gradient a Gradient?

Chris Nota, Philip S. Thomas

The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the di…

cs.LG2019

Reinforcement Learning When All Actions are Not Always Available

Yash Chandak, Georgios Theocharous, Blossom Metevier +1

The Markov decision process (MDP) formulation used to model many real-world sequential decision making problems does not efficiently capture the setting where the set of available…

cs.LG2019

Lifelong Learning with a Changing Action Set

Yash Chandak, Georgios Theocharous, Chris Nota +1

In many real-world sequential decision making problems, the number of available actions (decisions) can vary over time. While problems like catastrophic forgetting, changing transi…

cs.LG2019★ 11 cited

A Meta-MDP Approach to Exploration for Lifelong Reinforcement Learning

Francisco M. Garcia, Philip S. Thomas

In this paper we consider the problem of how a reinforcement learning agent that is tasked with solving a sequence of reinforcement learning problems (a sequence of Markov decision…