2 papers
cs.AI2025
Data-Efficient Safe Policy Improvement Using Parametric Structure
Kasper Engelen, Guillermo A. Pérez, Marnix Suilen
Safe policy improvement (SPI) is an offline reinforcement learning problem in which a new policy that reliably outperforms the behavior policy with high confidence needs to be comp…
cs.LO2025
Analyzing Value Functions of States in Parametric Markov Chains
Kasper Engelen, Guillermo A. Pérez, Shrisha Rao
Parametric Markov chains (pMC) are used to model probabilistic systems with unknown or partially known probabilities. Although (universal) pMC verification for reachability propert…