4 papers · 1 filter
Pruning the Path to Optimal Care: Identifying Systematically Suboptimal Medical Decision-Making with Inverse Reinforcement Learning
Inko Bovenzi, Adi Carmel, Michael Hu +5
In aims to uncover insights into medical decision-making embedded within observational data from clinical settings, we present a novel application of Inverse Reinforcement Learning…
Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
Leo Benac, Abhishek Sharma, Sonali Parbhoo +1
We consider the problem of estimating the transition dynamics from near-optimal expert trajectories in the context of offline model-based reinforcement learning. We develop a…
Decision-Point Guided Safe Policy Improvement
Abhishek Sharma, Leo Benac, Sonali Parbhoo +1
Within batch reinforcement learning, safe policy improvement (SPI) seeks to ensure that the learnt policy performs at least as well as the behavior policy that generated the datase…
Risk averse non-stationary multi-armed bandits
Leo Benac, Frédéric Godin
This paper tackles the risk averse multi-armed bandits problem when incurred losses are non-stationary. The conditional value-at-risk (CVaR) is used as the objective function. Two…