Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Computing the Reachability Value of Posterior-Deterministic POMDPs
Nathanaël Fijalkow, Arka Ghosh, Roman Kniazev +2
Partially observable Markov decision processes (POMDPs) are a fundamental model for sequential decision-making under uncertainty. However, many verification and synthesis problems…
cs.AI2025
Data-Efficient Safe Policy Improvement Using Parametric Structure
Kasper Engelen, Guillermo A. Pérez, Marnix Suilen
Safe policy improvement (SPI) is an offline reinforcement learning problem in which a new policy that reliably outperforms the behavior policy with high confidence needs to be comp…