538 citations · 1.8k across the 57 of their papers we have counts for
5 papers · 1 filter
Optimal policy evaluation using kernel-based temporal difference methods
Yaqi Duan, Mengdi Wang, Martin J. Wainwright
We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized…
Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
Andrea Zanette, Martin J. Wainwright, Emma Brunskill
Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically. We propose a new offline actor-critic algorithm that…
Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning
Koulik Khamaru, Eric Xia, Martin J. Wainwright +1
Various algorithms in reinforcement learning exhibit dramatic variability in their convergence rates and ultimate accuracy as a function of the problem structure. Such instance-spe…
Preference learning along multiple criteria: A game-theoretic perspective
Kush Bhatia, Ashwin Pananjady, Peter L. Bartlett +2
The literature on ranking from ordinal data is vast, and there are several ways to aggregate overall preferences from pairwise comparisons between objects. In particular, it is wel…
Minimax Off-Policy Evaluation for Multi-Armed Bandits
Cong Ma, Banghua Zhu, Jiantao Jiao +1
We study the problem of off-policy evaluation in the multi-armed bandit model with bounded rewards, and develop minimax rate-optimal procedures under three settings. First, when th…