1 paper
S. R. Eshwar, Aniruddha Mukherjee, Kintan Saha +4
In a recent work, we proposed Reliable Policy Iteration (RPI), that restores policy iteration's monotonicity-of-value-estimates property to the function approximation setting. Here…