Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Adaptive Policy Portfolios for Robust Markov Decision Processes
Kasper Engelen, Sebastian Junges, Guillermo A. Pérez +1
Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partia…
cs.AI2025
Data-Efficient Safe Policy Improvement Using Parametric Structure
Kasper Engelen, Guillermo A. Pérez, Marnix Suilen
Safe policy improvement (SPI) is an offline reinforcement learning problem in which a new policy that reliably outperforms the behavior policy with high confidence needs to be comp…