activity
20182022
most citedOff-Policy Policy Gradient with State Distribution Correction

44 citations · 79 across the 4 of their papers we have counts for

collaborators

10 papers

cs.LG2022

Provably Sample-Efficient RL with Side Information about Latent Dynamics

Yao Liu, Dipendra Misra, Miro Dudík +1

We study reinforcement learning (RL) in settings where observations are high-dimensional, but where an RL agent has access to abstract knowledge about the structure of the state sp…

cs.LG202035 cited

Provably Good Batch Reinforcement Learning Without Great Exploration

Yao Liu, Adith Swaminathan, Alekh Agarwal +1

Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is chall…

cs.LG2020

Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential Transitions

Omer Gottesman, Joseph Futoma, Yao Liu +4

Off-policy evaluation in reinforcement learning offers the chance of using observational data to improve future outcomes in domains such as healthcare and education, but safe deplo…

cs.LG2019

All-Action Policy Gradient Methods: A Numerical Integration Approach

Benjamin Petit, Loren Amdahl-Culleton, Yao Liu +2

While often stated as an instance of the likelihood ratio trick [Rubinstein, 1989], the original policy gradient theorem [Sutton, 1999] involves an integral over the action space.…

cs.LG2019

Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance Sampling

Yao Liu, Pierre-Luc Bacon, Emma Brunskill

Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods t…

cs.LG2019

Combining Parametric and Nonparametric Models for Off-Policy Evaluation

Omer Gottesman, Yao Liu, Scott Sussex +2

We consider a model-based approach to perform batch off-policy evaluation in reinforcement learning. Our method takes a mixture-of-experts approach to combine parametric and non-pa…