43 citations · 64 across the 26 of their papers we have counts for
6 papers · 1 filter
Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners
David Milec, Spyridon Samothrakis, Michael Fairbank +1
The dominance of Neural Networks (NNs) in RL is partially due to their incremental learning capability, which naturally suits the online, non-stationary nature of self-play trainin…
Best Agent Identification for General Game Playing
Matthew Stephenson, Alex Newcombe, Eric Piette +1
We present an efficient and generalised procedure to accurately identify the best (or near best) performing algorithm for each sub-task in a multi-problem domain. Our approach trea…
Anytime Sequential Halving in Monte-Carlo Tree Search
Dominic Sagers, Mark H. M. Winands, Dennis J. N. J. Soemers
Monte-Carlo Tree Search (MCTS) typically uses multi-armed bandit (MAB) strategies designed to minimize cumulative regret, such as UCB1, as its selection strategy. However, in the r…
Transfer of Fully Convolutional Policy-Value Networks Between Games and Game Variants
Dennis J. N. J. Soemers, Vegard Mella, Eric Piette +3
In this paper, we use fully convolutional architectures in AlphaZero-like self-play training setups to facilitate transfer between variants of board games as well as distinct games…
Manipulating the Distributions of Experience used for Self-Play Learning in Expert Iteration
Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson +1
Expert Iteration (ExIt) is an effective framework for learning game-playing policies from self-play. ExIt involves training a policy to mimic the search behaviour of a tree search…
Learning Policies from Self-Play with Policy Gradients and MCTS Value Estimates
Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson +1
In recent years, state-of-the-art game-playing agents often involve policies that are trained in self-playing processes where Monte Carlo tree search (MCTS) algorithms and trained…