works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.LG2026

Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs

Max Weltevrede, Caroline Horsch, Matthijs T. J. Spaan +1

The paper shows that training reinforcement‑learning agents on additional, irrelevant states acts like data augmentation and can improve zero‑shot generalization in contextual MDPs…

cs.LG2026

Generalization in offline RL: The structure is more important than the amount of pessimism

Max Weltevrede, Matthijs T. J. Spaan, Wendelin Böhmer

While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with hindering certain forms of generalization…

cs.AI2026

Shared Modular Recurrence in Contextual MDPs for Universal Morphology Control

Laurens Engwegen, Max Weltevrede, Caroline Horsch +2

A universal controller for any robot morphology would greatly improve computational and data efficiency. Steps have been made towards such multi-robot control by utilizing contextu…

cs.LG2026

Sparse Masked Attention Policies for Reliable Generalization

Caroline Horsch, Laurens Engwegen, Max Weltevrede +2

In reinforcement learning, abstraction methods that remove unnecessary information from the observation are commonly used to learn policies which generalize better to unseen tasks.…

cs.LG2025

How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning

Max Weltevrede, Moritz A. Zanger, Matthijs T. J. Spaan +1

In the zero-shot policy transfer setting in reinforcement learning, the goal is to train an agent on a fixed set of training environments so that it can generalise to similar, but…

cs.LG2025

Universal Value-Function Uncertainties

Moritz A. Zanger, Max Weltevrede, Yaniv Oren +4

Estimating epistemic uncertainty in value functions is a crucial challenge for many aspects of reinforcement learning (RL), including efficient exploration, safe decision-making, a…