66 citations · 104 across the 10 of their papers we have counts for
3 papers · 1 filter
Reinforcement Learning with Trajectory Feedback
Yonathan Efroni, Nadav Merlis, Shie Mannor
The standard feedback model of reinforcement learning requires revealing the reward of every visited state-action pair. However, in practice, it is often the case that such frequen…
Exploration-Exploitation in Constrained MDPs
Yonathan Efroni, Shie Mannor, Matteo Pirotta
In many sequential decision-making problems, the goal is to optimize a utility function while satisfying a set of constraints on different utilities. This learning problem is forma…
Optimistic Policy Optimization with Bandit Feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg +1
Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimizat…