activity
20182020
most citedAction Robust Reinforcement Learning and Applications in Continuous Control

66 citations · 96 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG202030 cited

Exploration-Exploitation in Constrained MDPs

Yonathan Efroni, Shie Mannor, Matteo Pirotta

In many sequential decision-making problems, the goal is to optimize a utility function while satisfying a set of constraints on different utilities. This learning problem is forma…

cs.LG2020

Optimistic Policy Optimization with Bandit Feedback

Yonathan Efroni, Lior Shani, Aviv Rosenberg +1

Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimizat…

cs.LG2019

Multi-step Greedy Reinforcement Learning Algorithms

Manan Tomar, Yonathan Efroni, Mohammad Ghavamzadeh

Multi-step greedy policies have been extensively used in model-based reinforcement learning (RL), both when a model of the environment is available (e.g.,~in the game of Go) and wh…

cs.LG2019

Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs

Lior Shani, Yonathan Efroni, Shie Mannor

Trust region policy optimization (TRPO) is a popular and empirically successful policy search algorithm in Reinforcement Learning (RL) in which a surrogate problem, that restricts…

cs.LG2019

Tight Regret Bounds for Model-Based Reinforcement Learning with Greedy Policies

Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh +1

State-of-the-art efficient model-based Reinforcement Learning (RL) algorithms typically act by iteratively solving empirical models, i.e., by performing \emph{full-planning} on Mar…

cs.LG201966 cited

Action Robust Reinforcement Learning and Applications in Continuous Control

Chen Tessler, Yonathan Efroni, Shie Mannor

A policy is said to be robust if it maximizes the reward while considering a bad, or even adversarial, model. In this work we formalize two new criteria of robustness to action unc…