66 citations · 104 across the 8 of their papers we have counts for
5 papers · 1 filter
Multi-step Greedy Reinforcement Learning Algorithms
Manan Tomar, Yonathan Efroni, Mohammad Ghavamzadeh
Multi-step greedy policies have been extensively used in model-based reinforcement learning (RL), both when a model of the environment is available (e.g.,~in the game of Go) and wh…
Online Planning with Lookahead Policies
Yonathan Efroni, Mohammad Ghavamzadeh, Shie Mannor
Real Time Dynamic Programming (RTDP) is an online algorithm based on Dynamic Programming (DP) that acts by 1-step greedy planning. Unlike DP, RTDP does not require access to the en…
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs
Lior Shani, Yonathan Efroni, Shie Mannor
Trust region policy optimization (TRPO) is a popular and empirically successful policy search algorithm in Reinforcement Learning (RL) in which a surrogate problem, that restricts…
Tight Regret Bounds for Model-Based Reinforcement Learning with Greedy Policies
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh +1
State-of-the-art efficient model-based Reinforcement Learning (RL) algorithms typically act by iteratively solving empirical models, i.e., by performing \emph{full-planning} on Mar…
Action Robust Reinforcement Learning and Applications in Continuous Control
Chen Tessler, Yonathan Efroni, Shie Mannor
A policy is said to be robust if it maximizes the reward while considering a bad, or even adversarial, model. In this work we formalize two new criteria of robustness to action unc…