7 papers
COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection
Guopeng Li, Moritz A. Zanger, Matthijs T. J. Spaan +1
Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-values are commonly learned by sep…
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
Moritz A. Zanger, Yijun Wu, Pascal R. Van der Vaart +2
Uncertainty quantification is central to safe and efficient deployments of deep learning models, yet many computationally practical methods lack lacking rigorous theoretical motiva…
Value Improved Actor Critic Algorithms
Yaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart +3
To learn approximately optimal acting policies for decision problems, modern Actor Critic algorithms rely on deep Neural Networks (DNNs) to parameterize the acting policy and greed…
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
Max Weltevrede, Moritz A. Zanger, Matthijs T. J. Spaan +1
In the zero-shot policy transfer setting in reinforcement learning, the goal is to train an agent on a fixed set of training environments so that it can generalise to similar, but…
Universal Value-Function Uncertainties
Moritz A. Zanger, Max Weltevrede, Yaniv Oren +4
Estimating epistemic uncertainty in value functions is a crucial challenge for many aspects of reinforcement learning (RL), including efficient exploration, safe decision-making, a…
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
Moritz A. Zanger, Pascal R. Van der Vaart, Wendelin Böhmer +1
Uncertainty quantification is a critical aspect of reinforcement learning and deep learning, with numerous applications ranging from efficient exploration and stable offline reinfo…