collaborators

7 papers

cs.RO2026

COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection

Guopeng Li, Moritz A. Zanger, Matthijs T. J. Spaan +1

Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-values are commonly learned by sep…

cs.LG2026

On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference

Moritz A. Zanger, Yijun Wu, Pascal R. Van der Vaart +2

Uncertainty quantification is central to safe and efficient deployments of deep learning models, yet many computationally practical methods lack lacking rigorous theoretical motiva…

cs.LG2026

Value Improved Actor Critic Algorithms

Yaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart +3

To learn approximately optimal acting policies for decision problems, modern Actor Critic algorithms rely on deep Neural Networks (DNNs) to parameterize the acting policy and greed…

cs.LG2025

How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning

Max Weltevrede, Moritz A. Zanger, Matthijs T. J. Spaan +1

In the zero-shot policy transfer setting in reinforcement learning, the goal is to train an agent on a fixed set of training environments so that it can generalise to similar, but…

cs.LG2025

Universal Value-Function Uncertainties

Moritz A. Zanger, Max Weltevrede, Yaniv Oren +4

Estimating epistemic uncertainty in value functions is a crucial challenge for many aspects of reinforcement learning (RL), including efficient exploration, safe decision-making, a…

cs.LG2025

Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model

Moritz A. Zanger, Pascal R. Van der Vaart, Wendelin Böhmer +1

Uncertainty quantification is a critical aspect of reinforcement learning and deep learning, with numerous applications ranging from efficient exploration and stable offline reinfo…