45 citations · 215 across the 20 of their papers we have counts for
3 papers · 1 filter
Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
Andrea Zanette, Martin J. Wainwright, Emma Brunskill
Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically. We propose a new offline actor-critic algorithm that…
Design of Experiments for Stochastic Contextual Linear Bandits
Andrea Zanette, Kefan Dong, Jonathan Lee +1
In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, t…
Universal Off-Policy Evaluation
Yash Chandak, Scott Niekum, Bruno Castro da Silva +3
When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must of…