papers

Publications (15)

cs.LG2024

Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization

Uri Gadot, Esther Derman, Navdeep Kumar +3

In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most ad…

cs.LG2026

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

Aneri Muni, Vincent Taboga, Esther Derman +2

Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral object…

cs.LG2023

Acting in Delayed Environments with Non-Stationary Markov Policies

Esther Derman, Gal Dalal, Shie Mannor

The standard Markov Decision Process (MDP) formulation hinges on the assumption that an action is executed immediately after it was chosen. However, assuming it is often unrealisti…

cs.LG2023

Twice Regularized Markov Decision Processes: The Equivalence between Robustness and Regularization

Esther Derman, Yevgeniy Men, Matthieu Geist +1

Robust Markov decision processes (MDPs) aim to handle changing or partially known system dynamics. To solve them, one typically resorts to robust optimization methods. However, thi…

cs.LG2026

Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism

Tianwei Ni, Esther Derman, Vineet Jain +3

Popular offline reinforcement learning (RL) methods rely on explicit conservatism, penalizing out-of-dataset actions or restricting rollout horizons. We question the universality o…

cs.LG2025

State Entropy Regularization for Robust Reinforcement Learning

Yonatan Ashlag, Uri Koren, Mirco Mutti +3

State entropy regularization has empirically shown better exploration and sample complexity in reinforcement learning (RL). However, its theoretical guarantees have not been studie…

cs.AI2024

Tree Search-Based Policy Optimization under Stochastic Execution Delay

David Valensi, Esther Derman, Shie Mannor +1

The standard formulation of Markov decision processes (MDPs) assumes that the agent's decisions are executed immediately. However, in numerous realistic applications such as roboti…

cs.LG2018

Soft-Robust Actor-Critic Policy-Gradient

Esther Derman, Daniel J. Mankowitz, Timothy A. Mann +1

Robust Reinforcement Learning aims to derive optimal behavior that accounts for model uncertainty in dynamical systems. However, previous studies have shown that by considering the…

cs.LG2024

Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis

Jia Lin Hau, Erick Delage, Esther Derman +2

In Markov decision processes (MDPs), quantile risk measures such as Value-at-Risk are a standard metric for modeling RL agents' preferences for certain outcomes. This paper propose…

cs.LG2025

Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning

Marco Jiralerspong, Esther Derman, Danilo Vucetic +5

A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with d…

cs.LG2023

Policy Gradient for Rectangular Robust Markov Decision Processes

Navdeep Kumar, Esther Derman, Matthieu Geist +2

Policy gradient methods have become a standard for training reinforcement learning agents in a scalable and efficient manner. However, they do not account for transition uncertaint…

cs.LG2021

Twice regularized MDPs and the equivalence between robustness and regularization

Esther Derman, Matthieu Geist, Shie Mannor

Robust Markov decision processes (MDPs) aim to handle changing or partially known system dynamics. To solve them, one typically resorts to robust optimization methods. However, thi…

cs.LG2019

A Bayesian Approach to Robust Reinforcement Learning

Esther Derman, Daniel Mankowitz, Timothy Mann +1

Robust Markov Decision Processes (RMDPs) intend to ensure robustness with respect to changing or adversarial system behavior. In this framework, transitions are modeled as arbitrar…

math.OC2020

Distributional Robustness and Regularization in Reinforcement Learning

Esther Derman, Shie Mannor

Distributionally Robust Optimization (DRO) has enabled to prove the equivalence between robustness and regularization in classification and regression, thus providing an analytical…

math.ST2017

Clustering and Model Selection via Penalized Likelihood for Different-sized Categorical Data Vectors

Esther Derman, Erwan Le Pennec

In this study, we consider unsupervised clustering of categorical vectors that can be of different size using mixture. We use likelihood maximization to estimate the parameters of…