activity
20122022
most citedPolicy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

45 citations · 215 across the 20 of their papers we have counts for

collaborators
Showing 2020Show all

8 papers · 1 filter

cs.LG20203 cited

Online Model Selection for Reinforcement Learning with Function Approximation

Jonathan N. Lee, Aldo Pacchiano, Vidya Muthukumar +2

Deep reinforcement learning has achieved impressive successes yet often requires a very large amount of interaction data. This result is perhaps unsurprising, as using complicated…

cs.LG2020

Provably Efficient Reward-Agnostic Navigation with Linear Value Iteration

Andrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer +1

There has been growing progress on theoretical analyses for provably efficient learning in MDPs with linear function approximation, but much of the existing work has made strong as…

cs.LG202035 cited

Provably Good Batch Reinforcement Learning Without Great Exploration

Yao Liu, Adith Swaminathan, Alekh Agarwal +1

Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is chall…

cs.LG20203 cited

Learning Abstract Models for Strategic Exploration and Fast Reward Transfer

Evan Zheran Liu, Ramtin Keramati, Sudarshan Seshadri +4

Model-based reinforcement learning (RL) is appealing because (i) it enables planning and thus more strategic exploration, and (ii) by decoupling dynamics from rewards, it enables f…

cs.AI2020

Value Driven Representation for Human-in-the-Loop Reinforcement Learning

Ramtin Keramati, Emma Brunskill

Interactive adaptive systems powered by Reinforcement Learning (RL) have many potential applications, such as intelligent tutoring systems. In such systems there is typically an ex…

stat.ML2020

Off-policy Policy Evaluation For Sequential Decisions Under Unobserved Confounding

Hongseok Namkoong, Ramtin Keramati, Steve Yadlowsky +1

When observed decisions depend only on observed features, off-policy policy evaluation (OPE) methods for sequential decision making problems can estimate the performance of evaluat…