activity
20122024
most citedFinite-Sample Analysis of Proximal Gradient TD Algorithms

101 citations · 225 across the 15 of their papers we have counts for

collaborators
Showing cs.LGShow all

17 papers · 1 filter

cs.LG2024

Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming

Xihong Su, Marek Petrik

Multi-model Markov decision process (MMDP) is a promising framework for computing policies that are robust to parameter uncertainty in MDPs. MMDPs aim to find a policy that maximiz…

cs.LG2024

Data Poisoning Attacks on Off-Policy Policy Evaluation Methods

Elita Lobo, Harvineet Singh, Marek Petrik +2

Off-policy Evaluation (OPE) methods are a crucial tool for evaluating policies in high-stakes domains such as healthcare, where exploration is often infeasible, unethical, or expen…

cs.LG2023

Bayesian Regret Minimization in Offline Bandits

Marek Petrik, Guy Tennenholtz, Mohammad Ghavamzadeh

We study how to make decisions that minimize Bayesian regret in offline linear bandits. Prior work suggests that one must take actions with maximum lower confidence bound (LCB) on…

cs.LG20211 cited

Policy Gradient Bayesian Robust Optimization for Imitation Learning

Zaynah Javed, Daniel S. Brown, Satvik Sharma +5

The difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are…

cs.LG20214 cited

Robust Maximum Entropy Behavior Cloning

Mostafa Hussein, Brendan Crowe, Marek Petrik +1

Imitation learning (IL) algorithms use expert demonstrations to learn a specific task. Most of the existing approaches assume that all expert demonstrations are reliable and trustw…

cs.LG2020

Soft-Robust Algorithms for Batch Reinforcement Learning

Elita A. Lobo, Mohammad Ghavamzadeh, Marek Petrik

In reinforcement learning, robust policies for high-stakes decision-making problems with limited data are usually computed by optimizing the percentile criterion, which minimizes t…