activity
20172021
most citedBeyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs

19 citations · 25 across the 8 of their papers we have counts for

collaborators

11 papers

cs.LG20212 cited

Lyapunov Robust Constrained-MDPs: Soft-Constrained Robustly Stable Policy Optimization under Model Uncertainty

Reazul Hasan Russel, Mouhacine Benosman, Jeroen Van Baar +1

Safety and robustness are two desired properties for any reinforcement learning algorithm. CMDPs can handle additional safety constraints and RMDPs can perform well under model unc…

cs.LG2020

Robust Constrained-MDPs: Soft-Constrained Robust Policy Optimization under Model Uncertainty

Reazul Hasan Russel, Mouhacine Benosman, Jeroen Van Baar

In this paper, we focus on the problem of robustifying reinforcement learning (RL) algorithms with respect to model uncertainties. Indeed, in the framework of model-based RL, we pr…

cs.LG2020

Entropic Risk Constrained Soft-Robust Policy Optimization

Reazul Hasan Russel, Bahram Behzadian, Marek Petrik

Having a perfect model to compute the optimal policy is often infeasible in reinforcement learning. It is important in high-stakes domains to quantify and manage risk induced by mo…

cs.LG20192 cited

Optimizing Norm-Bounded Weighted Ambiguity Sets for Robust MDPs

Reazul Hasan Russel, Bahram Behzadian, Marek Petrik

Optimal policies in Markov decision processes (MDPs) are very sensitive to model misspecification. This raises serious concerns about deploying them in high-stake domains. Robust M…

cs.AI2019

A Probabilistic Approach to Satisfiability of Propositional Logic Formulae

Reazul Hasan Russel

We propose a version of WalkSAT algorithm, named as BetaWalkSAT. This method uses probabilistic reasoning for biasing the starting state of the local search algorithm. Beta distrib…

cs.LG2019

Optimizing Percentile Criterion Using Robust MDPs

Bahram Behzadian, Reazul Hasan Russel, Marek Petrik +1

We address the problem of computing reliable policies in reinforcement learning problems with limited data. In particular, we compute policies that achieve good returns with high c…