Publications (15)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
Uri Gadot, Esther Derman, Navdeep Kumar +3
In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most ad…
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
Aneri Muni, Vincent Taboga, Esther Derman +2
Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral object…
Acting in Delayed Environments with Non-Stationary Markov Policies
Esther Derman, Gal Dalal, Shie Mannor
The standard Markov Decision Process (MDP) formulation hinges on the assumption that an action is executed immediately after it was chosen. However, assuming it is often unrealisti…
Twice Regularized Markov Decision Processes: The Equivalence between Robustness and Regularization
Esther Derman, Yevgeniy Men, Matthieu Geist +1
Robust Markov decision processes (MDPs) aim to handle changing or partially known system dynamics. To solve them, one typically resorts to robust optimization methods. However, thi…
Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
Tianwei Ni, Esther Derman, Vineet Jain +3
Popular offline reinforcement learning (RL) methods rely on explicit conservatism, penalizing out-of-dataset actions or restricting rollout horizons. We question the universality o…
State Entropy Regularization for Robust Reinforcement Learning
Yonatan Ashlag, Uri Koren, Mirco Mutti +3
State entropy regularization has empirically shown better exploration and sample complexity in reinforcement learning (RL). However, its theoretical guarantees have not been studie…
Tree Search-Based Policy Optimization under Stochastic Execution Delay
David Valensi, Esther Derman, Shie Mannor +1
The standard formulation of Markov decision processes (MDPs) assumes that the agent's decisions are executed immediately. However, in numerous realistic applications such as roboti…
Soft-Robust Actor-Critic Policy-Gradient
Esther Derman, Daniel J. Mankowitz, Timothy A. Mann +1
Robust Reinforcement Learning aims to derive optimal behavior that accounts for model uncertainty in dynamical systems. However, previous studies have shown that by considering the…
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
Jia Lin Hau, Erick Delage, Esther Derman +2
In Markov decision processes (MDPs), quantile risk measures such as Value-at-Risk are a standard metric for modeling RL agents' preferences for certain outcomes. This paper propose…
Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning
Marco Jiralerspong, Esther Derman, Danilo Vucetic +5
A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with d…
Policy Gradient for Rectangular Robust Markov Decision Processes
Navdeep Kumar, Esther Derman, Matthieu Geist +2
Policy gradient methods have become a standard for training reinforcement learning agents in a scalable and efficient manner. However, they do not account for transition uncertaint…
Twice regularized MDPs and the equivalence between robustness and regularization
Esther Derman, Matthieu Geist, Shie Mannor
Robust Markov decision processes (MDPs) aim to handle changing or partially known system dynamics. To solve them, one typically resorts to robust optimization methods. However, thi…
A Bayesian Approach to Robust Reinforcement Learning
Esther Derman, Daniel Mankowitz, Timothy Mann +1
Robust Markov Decision Processes (RMDPs) intend to ensure robustness with respect to changing or adversarial system behavior. In this framework, transitions are modeled as arbitrar…
Distributional Robustness and Regularization in Reinforcement Learning
Esther Derman, Shie Mannor
Distributionally Robust Optimization (DRO) has enabled to prove the equivalence between robustness and regularization in classification and regression, thus providing an analytical…
Clustering and Model Selection via Penalized Likelihood for Different-sized Categorical Data Vectors
Esther Derman, Erwan Le Pennec
In this study, we consider unsupervised clustering of categorical vectors that can be of different size using mixture. We use likelihood maximization to estimate the parameters of…