activity
20232026
most citedFinite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning

1 citations · 3 across the 7 of their papers we have counts for

collaborators

7 papers

cs.LG2026

Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning

Astrid Horn Brorholt, Maris F. L. Galesloot, Nils Jansen +2

Probabilistic shielding is a technique for safe reinforcement learning (RL). Typically, a static observer -- called the shield -- constrains the learning agent's actions to those f…

cs.AI2026

Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs

Joshua Wendland, Markel Zubia, Roman Andriushchenko +6

We introduce missingness-MDPs (miss-MDPs), a novel subclass of partially observable Markov decision processes (POMDPs) that incorporates the theory of missing data. A miss-MDP is a…

cs.LG2026

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

Maris F. L. Galesloot, Thomas Rhemrev, Nils Jansen

In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide guarantees on the (1) performance…

cs.AI2026★ 1 cited

Finite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning

David Hudák, Maris F. L. Galesloot, Martin Tappler +3

Solving partially observable Markov decision processes (POMDPs) requires computing policies under imperfect state information. Despite recent advances, the scalability of existing…

cs.AI2025

Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs

Maris F. L. Galesloot, Roman Andriushchenko, Milan Češka +2

Partially observable Markov decision processes (POMDPs) model specific environments in sequential decision-making under uncertainty. Critically, optimal policies for POMDPs may not…

cs.AI2024★ 1 cited

Pessimistic Iterative Planning with RNNs for Robust POMDPs

Maris F. L. Galesloot, Marnix Suilen, Thiago D. Simão +4

Robust POMDPs extend classical POMDPs to incorporate model uncertainty using so-called uncertainty sets on the transition and observation functions, effectively defining ranges of…