activity
20042023
most citedHigh-dimensional Ising model selection using -regularized logistic regression

538 citations · 1.8k across the 57 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

stat.ML20214 cited

Optimal policy evaluation using kernel-based temporal difference methods

Yaqi Duan, Mengdi Wang, Martin J. Wainwright

We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized…

cs.LG202114 cited

Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning

Andrea Zanette, Martin J. Wainwright, Emma Brunskill

Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically. We propose a new offline actor-critic algorithm that…

stat.ML20212 cited

Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning

Koulik Khamaru, Eric Xia, Martin J. Wainwright +1

Various algorithms in reinforcement learning exhibit dramatic variability in their convergence rates and ultimate accuracy as a function of the problem structure. Such instance-spe…

cs.LG2021

Preference learning along multiple criteria: A game-theoretic perspective

Kush Bhatia, Ashwin Pananjady, Peter L. Bartlett +2

The literature on ranking from ordinal data is vast, and there are several ways to aggregate overall preferences from pairwise comparisons between objects. In particular, it is wel…

stat.ML2021

Minimax Off-Policy Evaluation for Multi-Armed Bandits

Cong Ma, Banghua Zhu, Jiantao Jiao +1

We study the problem of off-policy evaluation in the multi-armed bandit model with bounded rewards, and develop minimax rate-optimal procedures under three settings. First, when th…