5 papers
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
Yaniv Oren, Viliam Vadocz, Joery A. de Vries +3
Monte Carlo Tree Search (MCTS) is a widely used approach for policy improvement through search with increasing popularity for real world applications. Due to the sequential and det…
Twice Sequential Monte Carlo for Tree Search
Yaniv Oren, Joery A. de Vries, Pascal R. van der Vaart +2
Model-based reinforcement learning (RL) methods that leverage search are responsible for many milestone breakthroughs in RL. Sequential Monte Carlo (SMC) recently emerged as an alt…
VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning
Joery A. de Vries, Jinke He, Yaniv Oren +3
Optimally trading-off exploration and exploitation is the holy grail of reinforcement learning as it promises maximal data-efficiency for solving any task. Bayes-optimal agents ach…
Trust-Region Twisted Policy Improvement
Joery A. de Vries, Jinke He, Yaniv Oren +1
Monte-Carlo tree search (MCTS) has driven many recent breakthroughs in deep reinforcement learning (RL). However, scaling MCTS to parallel compute has proven challenging in practic…
Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
Joery A. de Vries, Jinke He, Mathijs M. de Weerdt +1
Meta-reinforcement learning trains a single reinforcement learning agent on a distribution of tasks to quickly generalize to new tasks outside of the training set at test time. Fro…