5 papers
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
Joshua Wendland, Markel Zubia, Roman Andriushchenko +6
We introduce missingness-MDPs (miss-MDPs), a novel subclass of partially observable Markov decision processes (POMDPs) that incorporates the theory of missing data. A miss-MDP is a…
Robust Probabilistic Shielding for Safe Offline Reinforcement Learning
Maris F. L. Galesloot, Thomas Rhemrev, Nils Jansen
In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide guarantees on the (1) performance…
Finite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning
David Hudák, Maris F. L. Galesloot, Martin Tappler +3
Solving partially observable Markov decision processes (POMDPs) requires computing policies under imperfect state information. Despite recent advances, the scalability of existing…
Pessimistic Iterative Planning with RNNs for Robust POMDPs
Maris F. L. Galesloot, Marnix Suilen, Thiago D. Simão +4
Robust POMDPs extend classical POMDPs to incorporate model uncertainty using so-called uncertainty sets on the transition and observation functions, effectively defining ranges of…
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
Maris F. L. Galesloot, Roman Andriushchenko, Milan ÄeÅ¡ka +2
Partially observable Markov decision processes (POMDPs) model specific environments in sequential decision-making under uncertainty. Critically, optimal policies for POMDPs may not…