3 citations · 3 across the 1 of their papers we have counts for
4 papers
More for Less: Safe Policy Improvement With Stronger Performance Guarantees
Patrick Wienhöft, Marnix Suilen, Thiago D. Simão +3
In an offline reinforcement learning setting, the safe policy improvement (SPI) problem aims to improve the performance of a behavior policy according to which sample data has been…
Act-Then-Measure: Reinforcement Learning for Partially Observable Environments with Active Measuring
Merlijn Krale, Thiago D. Simão, Nils Jansen
We study Markov decision processes (MDPs), where agents have direct control over when and how they gather information, as formalized by action-contingent noiselessly observable MDP…
Decision-Making Under Uncertainty: Beyond Probabilities
Thom Badings, Thiago D. Simão, Marnix Suilen +1
This position paper reflects on the state-of-the-art in decision-making under uncertainty. A classical assumption is that probabilities can sufficiently capture all uncertainty in…
Safe Policy Improvement for POMDPs via Finite-State Controllers
Thiago D. Simão, Marnix Suilen, Nils Jansen
We study safe policy improvement (SPI) for partially observable Markov decision processes (POMDPs). SPI is an offline reinforcement learning (RL) problem that assumes access to (1)…