1 citations · 1 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023
More for Less: Safe Policy Improvement With Stronger Performance Guarantees
Patrick Wienhöft, Marnix Suilen, Thiago D. Simão +3
In an offline reinforcement learning setting, the safe policy improvement (SPI) problem aims to improve the performance of a behavior policy according to which sample data has been…
cs.LG2023
Strategy Synthesis in Markov Decision Processes Under Limited Sampling Access
Christel Baier, Clemens Dubslaff, Patrick Wienhöft +1
A central task in control theory, artificial intelligence, and formal methods is to synthesize reward-maximizing strategies for agents that operate in partially unknown environment…