activity
20122017
most citedPolicy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

45 citations · 85 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CL201723 cited

Generalized Grounding Graphs: A Probabilistic Framework for Understanding Grounded Commands

Thomas Kollar, Stefanie Tellex, Matthew Walter +8

Many task domains require robots to interpret and act upon natural language commands which are given by people and which refer to the robot's physical surroundings. Such interpreta…

cs.AI201710 cited

On Ensuring that Intelligent Machines Are Well-Behaved

Philip S. Thomas, Bruno Castro da Silva, Andrew G. Barto +1

Machine learning algorithms are everywhere, ranging from simple data analysis and pattern recognition tools used across the sciences to complex systems that achieve super-human per…

cs.AI201745 cited

Policy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

Philip S. Thomas, Emma Brunskill

We show how an action-dependent baseline can be used by the policy gradient theorem using function approximation, originally presented with action-independent baselines by (Sutton…

cs.AI20171 cited

Decoupling Learning Rules from Representations

Philip S. Thomas, Christoph Dann, Emma Brunskill

In the artificial intelligence field, learning often corresponds to changing the parameters of a parameterized function. A learning rule is an algorithm or mathematical expression…

cs.LG20173 cited

Sample Efficient Feature Selection for Factored MDPs

Zhaohan Daniel Guo, Emma Brunskill

In reinforcement learning, the state of the real world is often represented by feature vectors. However, not all of the features may be pertinent for solving the current task. We p…

cs.LG2016

A PAC RL Algorithm for Episodic POMDPs

Zhaohan Daniel Guo, Shayan Doroudi, Emma Brunskill

Many interesting real world domains involve reinforcement learning (RL) in partially observable environments. Efficient learning in such domains is important, but existing sample c…