5 citations · 11 across the 11 of their papers we have counts for
Showing 2020Show all
2 papers · 1 filter
cs.LG2020
Mirror Descent Policy Optimization
Manan Tomar, Lior Shani, Yonathan Efroni +1
Mirror descent (MD), a well-known first-order method in constrained convex optimization, has recently been shown as an important tool to analyze trust-region algorithms in reinforc…
cs.LG2020
Optimistic Policy Optimization with Bandit Feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg +1
Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimizat…