159 citations · 619 across the 30 of their papers we have counts for
3 papers · 1 filter
Active Model Estimation in Markov Decision Processes
Jean Tarbouriech, Shubhanshu Shekhar, Matteo Pirotta +2
We study the problem of efficient exploration in order to learn an accurate model of an environment, modeled as a Markov decision process (MDP). Efficient exploration in this probl…
Adaptive Sampling for Estimating Multiple Probability Distributions
Shubhanshu Shekhar, Tara Javidi, Mohammad Ghavamzadeh
We consider the problem of allocating samples to a finite set of discrete distributions in order to learn them uniformly well in terms of four common distance measures: ,…
Safe Policy Improvement by Minimizing Robust Baseline Regret
Marek Petrik, Yinlam Chow, Mohammad Ghavamzadeh
An important problem in sequential decision-making under uncertainty is to use limited data to compute a safe policy, i.e., a policy that is guaranteed to perform at least as well…