1 paper
Sarah Rathnam, Susan A. Murphy, Finale Doshi-Velez
In batch reinforcement learning, there can be poorly explored state-action pairs resulting in poorly learned, inaccurate models and poorly performing associated policies. Various r…