66 citations · 96 across the 2 of their papers we have counts for
7 papers
Exploration-Exploitation in Constrained MDPs
Yonathan Efroni, Shie Mannor, Matteo Pirotta
In many sequential decision-making problems, the goal is to optimize a utility function while satisfying a set of constraints on different utilities. This learning problem is forma…
Optimistic Policy Optimization with Bandit Feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg +1
Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimizat…
Multi-step Greedy Reinforcement Learning Algorithms
Manan Tomar, Yonathan Efroni, Mohammad Ghavamzadeh
Multi-step greedy policies have been extensively used in model-based reinforcement learning (RL), both when a model of the environment is available (e.g.,~in the game of Go) and wh…
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs
Lior Shani, Yonathan Efroni, Shie Mannor
Trust region policy optimization (TRPO) is a popular and empirically successful policy search algorithm in Reinforcement Learning (RL) in which a surrogate problem, that restricts…
Tight Regret Bounds for Model-Based Reinforcement Learning with Greedy Policies
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh +1
State-of-the-art efficient model-based Reinforcement Learning (RL) algorithms typically act by iteratively solving empirical models, i.e., by performing \emph{full-planning} on Mar…
Action Robust Reinforcement Learning and Applications in Continuous Control
Chen Tessler, Yonathan Efroni, Shie Mannor
A policy is said to be robust if it maximizes the reward while considering a bad, or even adversarial, model. In this work we formalize two new criteria of robustness to action unc…