activity
20162021
most citedPolicy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

45 citations · 138 across the 16 of their papers we have counts for

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2020

Learning Reusable Options for Multi-Task Reinforcement Learning

Francisco M. Garcia, Chris Nota, Philip S. Thomas

Reinforcement learning (RL) has become an increasingly active area of research in recent years. Although there are many algorithms that allow an agent to solve tasks efficiently, t…

cs.AI201710 cited

On Ensuring that Intelligent Machines Are Well-Behaved

Philip S. Thomas, Bruno Castro da Silva, Andrew G. Barto +1

Machine learning algorithms are everywhere, ranging from simple data analysis and pattern recognition tools used across the sciences to complex systems that achieve super-human per…

cs.AI201745 cited

Policy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

Philip S. Thomas, Emma Brunskill

We show how an action-dependent baseline can be used by the policy gradient theorem using function approximation, originally presented with action-independent baselines by (Sutton…

cs.AI20179 cited

Data-Efficient Policy Evaluation Through Behavior Policy Search

Josiah P. Hanna, Philip S. Thomas, Peter Stone +1

We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its per…

cs.AI20171 cited

Decoupling Learning Rules from Representations

Philip S. Thomas, Christoph Dann, Emma Brunskill

In the artificial intelligence field, learning often corresponds to changing the parameters of a parameterized function. A learning rule is an algorithm or mathematical expression…