4 papers
Optimistic Policy Optimization with Bandit Feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg +1
Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimizat…
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs
Lior Shani, Yonathan Efroni, Shie Mannor
Trust region policy optimization (TRPO) is a popular and empirically successful policy search algorithm in Reinforcement Learning (RL) in which a surrogate problem, that restricts…
Multi Instance Learning For Unbalanced Data
Mark Kozdoba, Edward Moroshko, Lior Shani +4
In the context of Multi Instance Learning, we analyze the Single Instance (SI) learning objective. We show that when the data is unbalanced and the family of classifiers is suffici…
Exploration Conscious Reinforcement Learning Revisited
Lior Shani, Yonathan Efroni, Shie Mannor
The Exploration-Exploitation tradeoff arises in Reinforcement Learning when one cannot tell if a policy is optimal. Then, there is a constant need to explore new actions instead of…