45 citations · 65 across the 4 of their papers we have counts for
5 papers
On Ensuring that Intelligent Machines Are Well-Behaved
Philip S. Thomas, Bruno Castro da Silva, Andrew G. Barto +1
Machine learning algorithms are everywhere, ranging from simple data analysis and pattern recognition tools used across the sciences to complex systems that achieve super-human per…
Policy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines
Philip S. Thomas, Emma Brunskill
We show how an action-dependent baseline can be used by the policy gradient theorem using function approximation, originally presented with action-independent baselines by (Sutton…
Data-Efficient Policy Evaluation Through Behavior Policy Search
Josiah P. Hanna, Philip S. Thomas, Peter Stone +1
We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its per…
Decoupling Learning Rules from Representations
Philip S. Thomas, Christoph Dann, Emma Brunskill
In the artificial intelligence field, learning often corresponds to changing the parameters of a parameterized function. A learning rule is an algorithm or mathematical expression…
Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning
Philip S. Thomas, Emma Brunskill
In this paper we present a new way of predicting the performance of a reinforcement learning policy given historical data that may have been generated by a different policy. The ab…