5 papers
CRAX: Fast Safe Reinforcement Learning Benchmarking
Tristan Tomilin, Mourad Boustani, Mickey Beurskens +1
Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progr…
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
Joshua Wendland, Markel Zubia, Roman Andriushchenko +6
We introduce missingness-MDPs (miss-MDPs), a novel subclass of partially observable Markov decision processes (POMDPs) that incorporates the theory of missing data. A miss-MDP is a…
Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response
Ariyan Bighashdel, Thiago D. Simão, Frans A. Oliehoek
Multi-agent reinforcement learning (MARL) offers a scalable alternative to exact game-theoretic analysis but suffers from non-stationarity and the need to maintain diverse populati…
Perception-Based Beliefs for POMDPs with Visual Observations
Miriam Schäfers, Merlijn Krale, Thiago D. Simão +2
Partially observable Markov decision processes (POMDPs) are a principled planning model for sequential decision-making under uncertainty. Yet, real-world problems with high-dimensi…
Tighter Value-Function Approximations for POMDPs
Merlijn Krale, Wietze Koops, Sebastian Junges +2
Solving partially observable Markov decision processes (POMDPs) typically requires reasoning about the values of exponentially many state beliefs. Towards practical performance, st…