20 citations · 66 across the 30 of their papers we have counts for
Showing 2021 · stat.MLShow all
2 papers · 2 filters
stat.ML2021
Accelerated and instance-optimal policy evaluation with linear function approximation
Tianjiao Li, Guanghui Lan, Ashwin Pananjady
We study the problem of policy evaluation with linear function approximation and present efficient and practical algorithms that come with strong optimality guarantees. We begin by…
stat.ML2021
Learning from an Exploring Demonstrator: Optimal Reward Estimation for Bandits
Wenshuo Guo, Kumar Krishna Agrawal, Aditya Grover +2
We introduce the "inverse bandit" problem of estimating the rewards of a multi-armed bandit instance from observing the learning process of a low-regret demonstrator. Existing appr…