4 papers
Twice Sequential Monte Carlo for Tree Search
Yaniv Oren, Joery A. de Vries, Pascal R. van der Vaart +2
Model-based reinforcement learning (RL) methods that leverage search are responsible for many milestone breakthroughs in RL. Sequential Monte Carlo (SMC) recently emerged as an alt…
Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
Isidoro Tamassia, Wendelin Böhmer
The AlphaZero framework provides a standard way of combining Monte Carlo planning with prior knowledge provided by a previously trained policy-value neural network. AlphaZero usual…
Shared Modular Recurrence in Contextual MDPs for Universal Morphology Control
Laurens Engwegen, Max Weltevrede, Caroline Horsch +2
A universal controller for any robot morphology would greatly improve computational and data efficiency. Steps have been made towards such multi-robot control by utilizing contextu…
Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs
Max Weltevrede, Caroline Horsch, Matthijs T. J. Spaan +1
In the zero-shot policy transfer (ZSPT) setting for contextual Markov decision processes (CMDP), agents train on a fixed, finite set of contexts and must generalize to new ones. Re…