1 paper
Andrew Cohen, Lei Yu, Robert Wright
We study an important yet under-addressed problem of quickly and safely improving policies in online reinforcement learning domains. As its solution, we propose a novel exploration…