3 papers
cs.AI2021
Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
Dylan Ashley, Anssi Kanervisto, Brendan Bennett
We present AlphaChute: a state-of-the-art algorithm that achieves superhuman performance in the ancient game of Chutes and Ladders. We prove that our algorithm converges to the Nas…
cs.LG2019
Incrementally Learning Functions of the Return
Brendan Bennett, Wesley Chung, Muhammad Zaheer +1
Temporal difference methods enable efficient estimation of value functions in reinforcement learning in an incremental fashion, and are of broader interest because they correspond…
cs.LG2018
Predicting Periodicity with Temporal Difference Learning
Kristopher De Asis, Brendan Bennett, Richard S. Sutton
Temporal difference (TD) learning is an important approach in reinforcement learning, as it combines ideas from dynamic programming and Monte Carlo methods in a way that allows for…