46 citations · 90 across the 5 of their papers we have counts for
5 papers
A Meta-MDP Approach to Exploration for Lifelong Reinforcement Learning
Francisco M. Garcia, Philip S. Thomas
In this paper we consider the problem of how a reinforcement learning agent that is tasked with solving a sequence of reinforcement learning problems (a sequence of Markov decision…
Learning Action Representations for Reinforcement Learning
Yash Chandak, Georgios Theocharous, James Kostas +2
Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the str…
Privacy Preserving Off-Policy Evaluation
Tengyang Xie, Philip S. Thomas, Gerome Miklau
Many reinforcement learning applications involve the use of data that is sensitive, such as medical records of patients or financial information. However, most current reinforcemen…
Data-Efficient Policy Evaluation Through Behavior Policy Search
Josiah P. Hanna, Philip S. Thomas, Peter Stone +1
We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its per…
Proximal Reinforcement Learning: A New Theory of Sequential Decision Making in Primal-Dual Spaces
Sridhar Mahadevan, Bo Liu, Philip Thomas +5
In this paper, we set forth a new vision of reinforcement learning developed by us over the past few years, one that yields mathematically rigorous solutions to longstanding import…