45 citations · 211 across the 18 of their papers we have counts for
12 papers · 1 filter
Sublinear Optimal Policy Value Estimation in Contextual Bandits
Weihao Kong, Gregory Valiant, Emma Brunskill
We study the problem of estimating the expected reward of the optimal policy in the stochastic disjoint linear bandit setting. We prove that for certain settings it is possible to…
Missingness as Stability: Understanding the Structure of Missingness in Longitudinal EHR data and its Impact on Reinforcement Learning in Healthcare
Scott L. Fleming, Kuhan Jeyapragasan, Tony Duan +4
There is an emerging trend in the reinforcement learning for healthcare literature. In order to prepare longitudinal, irregularly sampled, clinical datasets for reinforcement learn…
Problem Dependent Reinforcement Learning Bounds Which Can Identify Bandit Structure in MDPs
Andrea Zanette, Emma Brunskill
In order to make good decision under uncertainty an agent must learn from observations. To do so, two of the most common frameworks are Contextual Bandits and Markov Decision Proce…
Being Optimistic to Be Conservative: Quickly Learning a CVaR Policy
Ramtin Keramati, Christoph Dann, Alex Tamkin +1
While maximizing expected return is the goal in most reinforcement learning approaches, risk-sensitive objectives such as conditional value at risk (CVaR) are more suitable for man…
Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance Sampling
Yao Liu, Pierre-Luc Bacon, Emma Brunskill
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods t…
Directed Exploration for Reinforcement Learning
Zhaohan Daniel Guo, Emma Brunskill
Efficient exploration is necessary to achieve good sample efficiency for reinforcement learning in general. From small, tabular settings such as gridworlds to large, continuous and…