2 papers
cs.LG2024
Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes
David M. Bossens
The robust constrained Markov decision process (RCMDP) is a recent task-modelling framework for reinforcement learning that incorporates behavioural constraints and that provides r…
cs.LG2024
Low Variance Off-policy Evaluation with State-based Importance Sampling
David M. Bossens, Philip S. Thomas
In many domains, the exploration process of reinforcement learning will be too costly as it requires trying out suboptimal policies, resulting in a need for off-policy evaluation,…