activity
20122017
most citedSafe Policy Improvement by Minimizing Robust Baseline Regret

67 citations · 85 across the 4 of their papers we have counts for

collaborators

5 papers

cs.LG20171 cited

Value Directed Exploration in Multi-Armed Bandits with Structured Priors

Bence Cserna, Marek Petrik, Reazul Hasan Russel +1

Multi-armed bandits are a quintessential machine learning problem requiring the balancing of exploration and exploitation. While there has been progress in developing algorithms wi…

stat.ML201667 cited

Safe Policy Improvement by Minimizing Robust Baseline Regret

Marek Petrik, Yinlam Chow, Mohammad Ghavamzadeh

An important problem in sequential decision-making under uncertainty is to use limited data to compute a safe policy, i.e., a policy that is guaranteed to perform at least as well…

stat.ML2016

Building an Interpretable Recommender via Loss-Preserving Transformation

Amit Dhurandhar, Sechan Oh, Marek Petrik

We propose a method for building an interpretable recommender system for personalizing online content and promotions. Historical data available for the system consists of customer…

math.OC20154 cited

Robust Policy Optimization with Baseline Guarantees

Yinlam Chow, Marek Petrik, Mohammad Ghavamzadeh

Our goal is to compute a policy that guarantees improved return over a baseline policy even when the available MDP model is inaccurate. The inaccurate model may be constructed, for…

q-fin.PM201213 cited

An Approximate Solution Method for Large Risk-Averse Markov Decision Processes

Marek Petrik, Dharmashankar Subramanian

Stochastic domains often involve risk-averse decision makers. While recent work has focused on how to model risk in Markov decision processes using risk measures, it has not addres…