159 citations · 184 across the 7 of their papers we have counts for
8 papers
Visualizing MuZero Models
Joery A. de Vries, Ken S. Voskuil, Thomas M. Moerland +1
MuZero, a model-based reinforcement learning algorithm that uses a value equivalent dynamics model, achieved state-of-the-art performance in Chess, Shogi and the game of Go. In con…
The Second Type of Uncertainty in Monte Carlo Tree Search
Thomas M Moerland, Joost Broekens, Aske Plaat +1
Monte Carlo Tree Search (MCTS) efficiently balances exploration and exploitation in tree search based on count-derived uncertainty. However, these local visit counts ignore a secon…
Think Too Fast Nor Too Slow: The Computational Trade-off Between Planning And Reinforcement Learning
Thomas M. Moerland, Anna Deichler, Simone Baldi +2
Planning and reinforcement learning are two key approaches to sequential decision making. Multi-step approximate real-time dynamic programming, a recently successful algorithm clas…
The Potential of the Return Distribution for Exploration in RL
Thomas M. Moerland, Joost Broekens, Catholijn M. Jonker
This paper studies the potential of the return distribution for exploration in deterministic reinforcement learning (RL) environments. We study network losses and propagation mecha…
Monte Carlo Tree Search for Asymmetric Trees
Thomas M. Moerland, Joost Broekens, Aske Plaat +1
We present an extension of Monte Carlo Tree Search (MCTS) that strongly increases its efficiency for trees with asymmetry and/or loops. Asymmetric termination of search trees intro…
Efficient exploration with Double Uncertain Value Networks
Thomas M. Moerland, Joost Broekens, Catholijn M. Jonker
This paper studies directed exploration for reinforcement learning agents by tracking uncertainty about the value of each available action. We identify two sources of uncertainty t…