6 citations · 8 across the 22 of their papers we have counts for
1 paper · 2 filters
S. R. Eshwar, Aniruddha Mukherjee, Kintan Saha +4
In a recent work, we proposed Reliable Policy Iteration (RPI), that restores policy iteration's monotonicity-of-value-estimates property to the function approximation setting. Here…