Showing cs.AIShow all
2 papers · 1 filter
cs.AI2021
Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
Dylan Ashley, Anssi Kanervisto, Brendan Bennett
We present AlphaChute: a state-of-the-art algorithm that achieves superhuman performance in the ancient game of Chutes and Ladders. We prove that our algorithm converges to the Nas…
cs.AI2018
Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods
Craig Sherstan, Brendan Bennett, Kenny Young +4
This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function,…