157 citations · 349 across the 22 of their papers we have counts for
3 papers · 1 filter
On Minimax Optimal Offline Policy Evaluation
Lihong Li, Remi Munos, Csaba Szepesvari
This paper studies the off-policy evaluation problem, where one aims to estimate the value of a target policy based on a sample of observations collected by another policy. We firs…
Adaptive Monte Carlo via Bandit Allocation
James Neufeld, András György, Dale Schuurmans +1
We consider the problem of sequentially choosing between a set of unbiased Monte Carlo estimators to minimize the mean-squared-error (MSE) of a final combined estimate. By reducing…
Speeding Up Planning in Markov Decision Processes via Automatically Constructed Abstractions
Alejandro Isaza, Csaba Szepesvari, Vadim Bulitko +1
In this paper, we consider planning in stochastic shortest path (SSP) problems, a subclass of Markov Decision Problems (MDP). We focus on medium-size problems whose state space can…