1 paper
Andrew Patterson, Adam White, Martha White
Many reinforcement learning algorithms rely on value estimation, however, the most widely used algorithms -- namely temporal difference algorithms -- can diverge under both off-pol…